Dong Liang 0008

dblp:23/110-8 · DBLP profile ↗
← Back
69ranked-venue papers
17as first author
53since 2021 · last 2026
0000-0003-2784-3449ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 10 first-author · 29 since 2021Artificial intelligence and machine learning · 28 · 6 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA
abstract
Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical QA tasks, little attention has been paid to KE in multimodal medical scenarios. Unlike text-only settings, medical KE demands integrating updated knowledge with visual reasoning to support safe and interpretable clinical decisions. To address this gap, we propose MultiMedBench, the first benchmark tailored to evaluating KE in clinical multimodal tasks. Our framework spans both understanding and reasoning task types, defines a three-dimensional metric suite (reliability, generality, and locality), and supports cross-paradigm comparisons across general and domain-specific models. We conduct extensive experiments under single-editing and lifelong-editing settings. Results suggest that current methods struggle with generalization and long-tail reasoning, particularly in complex clinical workflows. We further present an efficiency analysis (e.g., edit latency, memory footprint), revealing practical trade-offs in real-world deployment across KE paradigms. Overall, MultiMedBench not only reveals the limitations of current approaches but also provides a solid foundation for developing clinically robust knowledge editing techniques in the future.
Shengtao Wen, Zhongying Pan, Xiang Chen 0016, Dong Liang 0008, Sheng-Jun Huang
AAAI8
2026 Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
Rong Quan, Yantao Lai, Dong Liang 0008, Jie Qin 0004
ICMR3
2026 Robust unsupervised visual tracking via image-to-video identity knowledge transferring
Bin Kang, Zongyu Wang, Dong Liang 0008, Tianyu Ding, Songlin Du
Pattern Recognit.3
2026 Robust Fine-Grained Visual Categorization via Cyclical Attention
abstract
Fine-grained visual categorization (FGVC) in open-world settings frequently encounters heavy occlusion (HO) samples that compromise discriminative features. However, effectively addressing heavy occlusion remains a challenge. Existing methods often either discard the occluded parts or utilize them through additional techniques such as image inpainting or multimodel strategies, each with its own set of advantages and limitations. In this article, we propose a novel approach inspired by human self-regulated learning (SRL) behavior: cyclical attention that leverages occluded regions through the attention recalibration in the feedback loop. In particular, we introduce a new multi-instance model where occluded parts are essential due to a special feedback structure at the basis of a cooperative game mechanism. This mimics SRL to re-evaluate the previous attention-based image patch selection strategy. We then embed the proposed multi-instance model into a transformer architecture, creating an SRL-FGVC transformer. The key innovation of this design is the cyclical attention, with the forward and feedback self-attention formulating a cooperative union to mitigate attention bias. Extensive experiments on six public datasets and an additional dataset we established demonstrate that the SRL-FGVC transformer consistently outperforms existing approaches in HO scenarios. This work presents a promising new direction for robust FGVC in challenging real-world conditions.
Bin Kang, Dong Liang 0008, Daoyuan Chen, Tianyu Ding, Mingqiang Wei
IEEE Trans. Neural Networks Learn. Syst.2
2026 Search by Image: Deeply Exploring Beneficial Features for Beauty Product Retrieval
abstract
Searching by image is popular yet still challenging in e-commerce due to the extensive interference arising from (i) data variations (e.g., background, pose, visual angle, brightness) of real-world captured images and (ii) similar images in the query dataset. This article studies a practically meaningful problem of beauty product retrieval (BPR) by neural networks. We broadly extract different types of image features and raise an intriguing question that whether these features are beneficial to (i) suppress data variations of real-world captured images and (ii) distinguish one image from others which look very similar but are intrinsically different beauty products in the dataset, therefore leading to an enhanced capability of BPR. To answer it, we present a novel v ariable-attention neural network to understand the combination of m ultiple features (termed VM-Net) of beauty product images. Considering that there are few publicly released training datasets for BPR, we establish a new dataset with more than one million images classified into more than 20K categories to improve both the generalization and anti-interference abilities of VM-Net and other methods. We verify the performance of VM-Net and its competitors on the benchmark dataset Perfect-500K, where VM-Net shows clear improvements over the competitors in terms of \(MAP@7\) . The source code and dataset will be released upon publication.
Mingqiang Wei, Haoran Xie 0001, Dong Liang 0008, Dingkun Zhu, Fu Lee Wang
ACM Trans. Multim. Comput. Commun. Appl.4
2026 Dual-Branch Aesthetic Image Retouching via Active Reinforcement Learning for Color Enhancement and Composition Optimization
abstract
Existing learning-based visual retouching primarily focuses on improving image quality through end-to-end objective mapping between input and retouched images. However, these approaches often overlook two critical aspects: the progressive nature of image retouching and the subjective aesthetic preferences, resulting in suboptimal visual outcomes. To address this, we introduce Automatic Aesthetic Image Retouching via active reinforcement learning (A$^{3}$3RL) to enhance the visualization experience in two sub-tasks: color enhancement and composition optimization, which are formulated as a unified Markov Decision Process in the proposed A$^{3}$3RL framework. In our approach, each pixel functions as an autonomous agent that determines optimal actions based on aesthetic guidance, engaging in online exploration through immediate pixel-wise and channel-wise feedback from the aesthetic environment. By leveraging a pretrained image aesthetic model, our method ensures that the A$^{3}$3RL process aligns with human aesthetic preferences and adheres to subjective aesthetic principles. The framework integrates pixel-level retouching actions with image-level operations to achieve optimal image sequences through progressive iterations. Extensive experiments demonstrate that our method effectively recalibrates image aesthetics across multiple dimensions: low-level quality metrics (PSNR, SSIM), visual perception (LPIPS), and subjective visual experience (human survey). The results demonstrate high consistency with expert-retouched ground-truth images.
Dong Liang 0008, Yuanhang Gao, Sheng-Jun Huang, Songcan Chen
IEEE Trans. Vis. Comput. Graph.1
2025 StructSR: Refuse Spurious Details in Real-World Image Super-Resolution
abstract
Diffusion-based models have shown great promise in real-world image super-resolution (Real-ISR), but often generate content with structural errors and spurious texture details due to the empirical priors and illusions of these models. To address this issue, we introduce StructSR, a simple, effective, and plug-and-play method that enhances structural fidelity and suppresses spurious details for diffusion-based Real-ISR. StructSR operates without the need for additional fine-tuning, external model priors, or high-level semantic knowledge. At its core is the Structure-Aware Screening (SAS) mechanism, which identifies the image with the highest structural similarity to the low-resolution (LR) input in the early inference stage, allowing us to leverage it as a historical structure knowledge to suppress the generation of spurious details. By intervening in the diffusion inference process, StructSR seamlessly integrates with existing diffusion-based Real-ISR models. Our experimental results demonstrate that StructSR significantly improves the fidelity of structure and texture, improving the PSNR and SSIM metrics by an average of 5.27% and 9.36% on a synthetic dataset (DIV2K-Val) and 4.13% and 8.64% on two real-world datasets (RealSR and DRealSR) when integrated with four state-of-the-art diffusion-based Real-ISR methods.
Dong Liang 0008, Tianyu Ding, Sheng-Jun Huang
AAAI2
2025 CLIPGaze: Zero-Shot Goal-Directed Scanpath Prediction Using CLIP
abstract
Goal-directed scanpath prediction aims to predict people’s gaze shift path when searching for objects in a visual scene. Most existing goal-directed scanpath prediction methods cannot generalize to target classes not present during training. Besides, they usually exploit different pre-trained models to extract features for the target prompt and image, resulting in big feature gap and making the subsequent feature matching and fusion very difficult. To solve the above problems, we propose a novel zero-shot goal-directed scanpath prediction model named CLIPGaze. We use CLIP to extract pre-matched features for the target prompt and input image, making the feature fusion easier to receive. Using large model like CLIP can also enhance the whole model’s generalization ability on target classes not present during training. We propose a hierarchical visual-semantic feature fusion module to fuse the target and image features more comprehensively. Furthermore, due to the limited number of classes in goal-directed scanpath dataset, we employ image segmentation as a proxy task to help train the feature fusion module, significantly enhancing our model’s performance in zeroshot setting. Extensive experiments demonstrate the effectiveness of our method on both seen and unseen target classes.
Yantao Lai, Rong Quan, Dong Liang 0008, Jie Qin 0004
ICASSP3
2025 InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention Prior
abstract
3D Gaussian Splatting based 3D editing has demonstrated impressive performance in recent years. However, the multi-view editing often exhibits significant local inconsistency, especially in areas of non-rigid deformation, which lead to local artifacts, texture blurring, or semantic variations in edited 3D scenes. We also found that the existing editing methods, which rely entirely on text prompts make the editing process a "one-shot deal", making it difficult for users to control the editing degree flexibly. In response to these challenges, we present InterGSEdit, a novel framework for high-quality 3DGS editing via interactively selecting key views with users' preferences. We propose a CLIP-based Semantic Consistency Selection (CSCS) strategy to adaptively screen a group of semantically consistent reference views for each user-selected key view. Then, the cross-attention maps derived from the reference views are used in a weighted Gaussian Splatting unprojection to construct the 3D Geometry-Consistent Attention Prior ($GAP^{3D}$). We project $GAP^{3D}$ to obtain 3D-constrained attention, which are fused with 2D cross-attention via Attention Fusion Network (AFN). AFN employs an adaptive attention strategy that prioritizes 3D-constrained attention for geometric consistency during early inference, and gradually prioritizes 2D cross-attention maps in diffusion for fine-grained features during the later inference. Extensive experiments demonstrate that InterGSEdit achieves state-of-the-art performance, delivering consistent, high-fidelity 3DGS editing with improved user experience.
Minghao Wen, Shengjie Wu, Kangkan Wang, Dong Liang 0008
ICCV4
2025 Equiangular Aligned Dual Prompt Learning for Open-Set Recognition
Enhao Zhang 0002, Dong Liang 0008, Chuanxing Geng
ICIC (9)3
2025 LoD: Loss-difference OOD Detection by Intentionally Label-Noisifying Unlabeled Wild Data
abstract
Using unlabeled wild data containing both in-distribution (ID) and out-of-distribution (OOD) data to improve the safety and reliability of models has recently received increasing attention. Existing methods either design customized losses for labeled ID and unlabeled wild data then perform joint optimization, or first filter out OOD data from the latter then learn an OOD detector. While achieving varying degrees of success, two potential issues remain: (i) Labeled ID data typically dominates the learning of models, inevitably making models tend to fit OOD data as IDs; (ii) The selection of thresholds for identifying OOD data in unlabeled wild data usually faces dilemma due to the unavailability of pure OOD samples. To address these issues, we propose a novel loss-difference OOD detection framework (LoD) by intentionally label-noisifying unlabeled wild data. Such operations not only enable labeled ID data and OOD data in unlabeled wild data to jointly dominate the models' learning but also ensure the distinguishability of the losses between ID and OOD samples in unlabeled wild data, allowing the classic clustering technique (e.g., K-means) to filter these OOD samples without requiring thresholds any longer. We also provide theoretical foundation for LoD's viability, and extensive experiments verify its superiority.
Chuanxing Geng, Xinrui Wang 0003, Dong Liang 0008, Songcan Chen, Pong C. Yuen
IJCAI4
2025 E²GO : Free Your Hands for Smartphone Interaction
abstract
Current eye-gaze interaction technologies for smartphones are considered inflexible, inaccurate, and power-hungry. These methods typically rely on hand involvement and accomplish partial interactions. In this paper, we propose a novel eye-gaze smartphone interaction method named Event-driven Eye-Gaze Operation (E2GO), which can realize comprehensive interaction using only eyes and gazes to cover various interaction types. Before the interaction, an anti-jitter gaze estimation method was exploited to stabilize human eye fixation and predict accurate and stable human gaze positions on smartphone screens to further explore refined time-dependent eye-gaze interactions. We also integrated an event-triggering mechanism in E2GO to significantly decrease its power consumption to deploy on smartphones. We have implemented the prototype of E2GO on different brands of smartphones and conducted a comprehensive user study to validate its efficacy, demonstrating E2GO‘s superior smartphone control capabilities across various scenarios. Demo videos
Shaoming Yan, Yuanliang Ju, Rong Quan, Huawei Tu, Dong Liang 0008
Int. J. Hum. Comput. Interact.5
2025 SemMatcher: Semantic-aware feature matching with neighborhood consensus
Qimin Jiang, Xiaoyong Lu, Dong Liang 0008, Songlin Du
J. Vis. Commun. Image Represent.3
2025 SGD-font: Style and glyph decoupling for one-shot font generation
Dong Liang 0008
Knowl. Based Syst.3
2025 Aesthetics-Guided Low-Light Enhancement
abstract
Evaluating the performance of low-light image enhancement (LLE) is highly subjective, thus making integrating human preferences into LLE a necessity. Existing methods fail to consider this and present a series of potentially valid heuristic criteria for training LLE models. In this paper, we propose a new paradigm, i.e., aesthetics-guided low-light image enhancement (ALL-E), which introduces aesthetic preferences to LLE and motivates training in a reinforcement learning framework with an aesthetic reward. Each pixel, functioning as an agent, refines itself by recursive actions. We further present ALL-E+, an extended version of ALL-E, which casts a two-stage aesthetics-guided enhancement and denoising. ALL-E+ achieves low-light enhancement and denoising compensation sequentially in a unified framework, resulting in significant improvements in both subjective visual experience and objective evaluation. Extensive experiments show that integrating aesthetic preferences can further improve the visual experience of enhanced images. Our results on various benchmarks also demonstrate the superiority of our method over state-of-the-art methods.
Dong Liang 0008, Yuanhang Gao, Ling Li 0010, Zhengyan Xu, Sheng-Jun Huang, Songcan Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Image Lens Flare Removal Using Adversarial Curve Learning
abstract
When taking images against strong light sources, the resulting images often contain heterogeneous flare artifacts. These artifacts can significantly affect image visual quality and downstream computer vision tasks. While collecting real data pairs of flare-corrupted/flare-free images for training flare removal models is challenging, current methods utilize the direct-add approach to synthesize training data. However, these methods do not consider automatic exposure and tone mapping in the image signal processing pipeline (ISP), leading to the limited generalization capability of deep model training using such data. Besides, existing light source recovery methods hardly recover multiple light sources due to the different sizes, shapes, and illuminance of various light sources. In this paper, we propose a solution to improve the performance of lens flare removal by revisiting the ISP, remodeling the principle of automatic exposure in the synthesis pipeline, and designing a more reliable light source recovery strategy. The new pipeline approaches realistic imaging by discriminating the local and global illumination through a convex combination, avoiding global illumination shifting and local over-saturation. Moreover, the current deep models are only generalized to specific devices due to the diversity of cameras' ISPs. To achieve better generalization on different devices, we formulate the generalization problem as an adversarial training problem and embed an adversarial curve learning (ACL) paradigm in the synthesis pipeline to gain better performance. For recovering multiple light sources, our strategy convexly averages the input and output of the neural network based on illuminance levels, thereby avoiding the need for a hard threshold in identifying light sources. We also contribute a new flare removal testing dataset containing the flare-corrupted images captured by fifteen types of consumer electronics. The dataset facilitates the verification of the generalization capability of flare removal methods. Extensive experiments show that our solution can effectively improve the performance of lens flare removal and push the frontier toward more general situations.
Yuyan Zhou, Dong Liang 0008, Songcan Chen, Sheng-Jun Huang
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 SSDQ: Target Speaker Extraction via Semantic and Spatial Dual Querying
abstract
Target Speaker Extraction (TSE) in real-world multi-speaker environments is highly challenging. Previous works have largely relied on pre-enrollment speech to extract the target speaker's voice. However, such methods are limited in spontaneous scenarios where pre-enrollment speech or spatial information is unavailable. To address this, we propose Semantic and Spatial Dual Querying (SSDQ), a unified framework that integrates natural language descriptions and region-based spatial queries to guide TSE. SSDQ employs dual query encoders for semantic and spatial cues, fusing them into the audio stream via a FiLM-based interaction module. A novel Controllable Feature Wrapping (CFW) mechanism further enables a dynamic balance between speaker identity and acoustic clarity. We also introduce SS-Libri, a spatialized mixture dataset designed to benchmark dual-query systems. Extensive experiments demonstrate that SSDQ achieves superior extraction accuracy and robustness under challenging conditions, yielding the SI-SNRi of 19.63 dB, SNRi of 20.30 dB, PESQ of 1.83, and STOI of 0.26.
Xinjia Zhu, Xinyuan Qian 0001, Dong Liang 0008
IEEE Signal Process. Lett.3
2025 Handling Noisy Annotation for Remote Sensing Semantic Segmentation via Boundary-Aware Knowledge Distillation
abstract
In recent years, image segmentation has made significant progress, but acquiring annotated data is still a considerable challenge, especially in remote sensing imagery (RSI). The complex structure and inter-category confusion of RSI increase the time-consuming and cost of pixel-level annotation, and noisy annotations inevitably appear. This paper proposes a boundary-aware knowledge distillation method (BAKD) to handle noisy annotations by evaluating their uncertainty. BAKD consists of two core strategies: Predictive Confidence Evaluation (PCE) and Boundary-annotated Reliability Evaluation (BRE). The predictive confidence jointly decided by the teacher and student networks reflects the annotation’s uncertainty. The boundary-annotated reliability directly measures the annotation’s uncertainty based on the distance from the annotation to the semantic boundary. Leveraging these two types of uncertainty information, BAKD assigns each sample a comprehensive boundary-aware weight to identify samples with potential noisy annotations. This alleviates the impact of noisy annotation on the model’s training and improves its generalization performance. Experimental results show that BAKD achieves competitive semantic segmentation performance on the Potsdam and Vaihingen benchmarks compared with the state-of-the-art KD methods. In addition, BAKD can be easily integrated into semantic segmentation methods based on KD, extending their applicability in handling noisy annotations. Codes are available at https://github.com/sunyueue/BAKD.git.
Dong Liang 0008, Shao-Yuan Li, Songcan Chen, Sheng-Jun Huang
IEEE Trans. Geosci. Remote. Sens.2
2025 Diffusion-Noise-Based Augmentation for Long-Tailed Remote Sensing Image Classification
abstract
Remote sensing image classification refers to the task that using algorithms to categorize satellite or aerial imagery into different land cover types. In the real world, long-tailed data distribution is commonly present in remote sensing image classification tasks, causing models to excessively favor sufficient head classes during training and degrading prediction accuracy for scarce tail classes. Although existing methods such as resampling and re-weighting can alleviate the issue of data imbalance to a certain extent, they struggle to sufficiently enhance the diversity of tail class samples. In recent years, some works have begun to use diffusion models to generate diverse samples to balance the class distribution. However, these approaches often overlook the distribution inconsistency between generated and real images, which will hinder the improvement of model performance. To tackle these challenges, this paper proposes a novel diffusion-noise-based augmentation method (DONA) with a two-stage training process. Before training, our specially designed conditional prompts are used together with the original training set to guide the diffusion model in image generation. Furthermore, we propose two strategies to effectively leverage the generated images, which are applied respectively at the end of the first training stage and during the second training stage. First, we design DiffCam-Mix to fuse the background of the generated data with the foreground of the original data, preserving the essential information of the original real images while incorporating the diversity of the generated ones. Second, we use cosine similarity to minimize the differences between the mixed data and their corresponding original data, further calibrating the distribution of different samples. Extensive experiments on three public datasets—SIRI-WHU-LT, PatternNet-LT, and RSI-CB256-LT—demonstrate the effectiveness of the proposed method.
Qianqian Wang 0014, Haibo Ye, Dong Liang 0008, Sheng-Jun Huang
IEEE Trans. Geosci. Remote. Sens.3
2025 CrossNet: Cross-Scene Background Subtraction Network via 3D Optical Flow
abstract
This paper investigates an intriguing yet unsolved problem of cross-scene background subtraction for training only one deep model to process large-scale video streaming. We propose an end-to-end cross-scene background subtraction network via 3D optical flow, dubbed CrossNet. First, we design a new motion descriptor, hierarchical 3D optical flows (3D-HOP), to observe fine-grained motion. Then, we build a cross-modal dynamic feature filter (CmDFF) to enable the motion and appearance feature interaction. CrossNet exhibits better generalization since the proposed modules are encouraged to learn more discriminative semantic information between the foreground and the background. Furthermore, we design a loss function to balance the size diversity of foreground instances since small objects are usually missed due to training bias. Our whole background subtraction model is called Hierarchical Optical Flow Attention Model (HOFAM). Unlike most of the existing stochastic-process-based and CNN-based background subtraction models, HOFAM will avoid inaccurate online model updating, not heavily rely on scene-specific information, and well represent ambient motion in the open world. Experimental results on several well-known benchmarks demonstrate that it outperforms state-of-the-art by a large margin. The proposed framework can be flexibly integrated into arbitrary streaming media systems in a plug-and-play form. Codes are available athttps://github.com/dongzhang89/HOFAM.
Dong Liang 0008, Qiong Wang 0001, Zongqi Wei, Liyan Zhang 0001
IEEE Trans. Multim.1
2024 Pathformer3D: A 3D Scanpath Transformer for $360^{\circ }$ Images
Rong Quan, Yantao Lai, Mengyu Qiu, Dong Liang 0008
ECCV (35)4
2024 Aesthetics-Driven Active Reinforcement Learning for Color Enhancement
Yuanhang Gao, Qi Zhu 0001, Yutian Fu, Dong Liang 0008
ICIC (11)4
2024 Beyond the Limits: Tackling Extreme Overexposure with Diffusion Model
Zhengyan Xu, Dong Liang 0008
ICIC (11)4
2024 Relative difficulty distillation for semantic segmentation
Dong Liang 0008, Songcan Chen, Sheng-Jun Huang
Sci. China Inf. Sci.1
2024 CPRNC: Channels pruning via reverse neuron crowding for model compression
Pingfan Wu, Hengyi Huang, Dong Liang 0008, Ningzhong Liu
Comput. Vis. Image Underst.4
2024 PIE: Physics-Inspired Low-Light Enhancement
Dong Liang 0008, Zhengyan Xu, Ling Li 0010, Mingqiang Wei, Songcan Chen
Int. J. Comput. Vis.1
2024 Fine-grained recognition via submodular optimization regulated progressive training
Bin Kang, Songlin Du, Dong Liang 0008, Xin Li 0086
Pattern Recognit.3
2024 eViTBins: Edge-Enhanced Vision-Transformer Bins for Monocular Depth Estimation on Edge Devices
abstract
Monocular depth estimation (MDE) remains a fundamental yet not well-solved problem in computer vision. Current wisdom of MDE often achieves blurred or even indistinct depth boundaries, degenerating the quality of vision-based intelligent transportation systems. This paper presents an edge-enhanced vision transformer bins network for monocular depth estimation, termed eViTBins. eViTBins has three core modules to predict monocular depth maps with exceptional smoothness, accuracy, and fidelity to scene structures and object edges. First, a multi-scale feature fusion module is proposed to circumvent the loss of depth information at various levels during depth regression. Second, an image-guided edge-enhancement module is proposed to accurately infer depth values around image boundaries. Third, a vision transformer-based depth discretization module is introduced to comprehend the global depth distribution. Meanwhile, unlike most MDE models that rely on high-performance GPUs, eViTBins is optimized for seamless deployment on edge devices, such as NVIDIA Jetson Nano and Google Coral SBC, making it ideal for real-time intelligent transportation systems applications. Extensive experimental evaluations corroborate the superiority of eViTBins over competing methods, notably in terms of preserving depth edges and global depth representations.
Yutong She, Peng Li 0064, Mingqiang Wei, Dong Liang 0008, Yiping Chen 0002, Haoran Xie 0001, Fu Lee Wang
IEEE Trans. Intell. Transp. Syst.4
2024 RT-less: a multi-scene RGB dataset for 6D pose estimation of reflective texture-less objects
Xinyue Zhao, Quanzhi Li, Yue Chao, Quanyou Wang, Zaixing He, Dong Liang 0008
Vis. Comput.6
2023 MSFORMER: Multi-Scale Transformer with Neighborhood Consensus for Feature Matching
abstract
Existing feature matching methods tend to extract feature descriptors by feeding down-sampled feature maps into a Transformer that is unable to extend feature scales, leading to false correspondences between small-size objects. This paper proposes MSFormer, which uses Transformers situated in different branches to obtain feature descriptors. In one branch, convolutions are integrated into self-attention layers elegantly to compensate for the lack of the local structure information. In another branch, a multi-scale Transformer is proposed through injecting heterogeneous receptive field sizes into tokens. Additionally, a neighborhood consensus mechanism is proposed by re-ranking initial matches to make a constraint of geometric consensus on neighborhood feature descriptors. Extensive experiments on indoor and outdoor pose estimations show that MSFormer outperforms existing state-of-the- art methods by a large margin.
Yaping Yan, Dong Liang 0008, Songlin Du
ICASSP3
2023 Improving Lens Flare Removal with General-Purpose Pipeline and Multiple Light Sources Recovery
abstract
When taking images against strong light sources, the resulting images often contain heterogeneous flare artifacts. These artifacts can importantly affect image visual quality and downstream computer vision tasks. While collecting real data pairs of flare-corrupted/flare-free images for training flare removal models is challenging, current methods utilize the direct-add approach to synthesize data. However, these methods do not consider automatic exposure and tone mapping in image signal processing pipeline (ISP), leading to the limited generalization capability of deep models training using such data. Besides, existing methods struggle to handle multiple light sources due to the different sizes, shapes and illuminance of various light sources. In this paper, we propose a solution to improve the performance of lens flare removal by revisiting the ISP and remodeling the principle of automatic exposure in the synthesis pipeline and design a more reliable light sources recovery strategy. The new pipeline approaches realistic imaging by discriminating the local and global illumination through convex combination, avoiding global illumination shifting and local over-saturation. Our strategy for recovering multiple light sources convexly averages the input and output of the neural network based on illuminance levels, thereby avoiding the need for a hard threshold in identifying light sources. We also contribute a new flare removal testing dataset containing the flare-corrupted images captured by ten types of consumer electronics. The dataset facilitates the verification of the generalization capability of flare removal methods. Extensive experiments show that our solution can effectively improve the performance of lens flare removal and push the frontier toward more general situations.
Yuyan Zhou, Dong Liang 0008, Songcan Chen, Sheng-Jun Huang, Chongyi Li
ICCV2
2023 ALL-E: Aesthetics-guided Low-light Image Enhancement
abstract
Evaluating the performance of low-light image enhancement (LLE) is highly subjective, thus making integrating human preferences into image enhancement a necessity. Existing methods fail to consider this and present a series of potentially valid heuristic criteria for training enhancement models. In this paper, we propose a new paradigm, i.e., aesthetics-guided low-light image enhancement (ALL-E), which introduces aesthetic preferences to LLE and motivates training in a reinforcement learning framework with an aesthetic reward. Each pixel, functioning as an agent, refines itself by recursive actions, i.e., its corresponding adjustment curve is estimated sequentially. Extensive experiments show that integrating aesthetic assessment improves both subjective experience and objective evaluation. Our results on various benchmarks demonstrate the superiority of ALL-E over state-of-the-art methods. Source code: https://dongl-group.github.io/project pages/ALLE.html
Ling Li 0010, Dong Liang 0008, Yuanhang Gao, Sheng-Jun Huang, Songcan Chen
IJCAI2
2023 Visual ScanPath Transformer: Guiding Computers to See the World
abstract
We propose to exploit the scanpath prediction technology to simulate human visual system to automatically generate gaze scanpaths for VR/AR applications, to alleviate the equipment and computational cost in foveated rendering. Specifically, we propose a novel deep learning-based scanpath prediction model called Visual ScanPath Transformer (VSPT), to predict human gaze scanpaths in both free viewing and task-driven viewing situations, based on which the VR/AR systems can execute foveated rendering rapidly and cheaply. The proposed VSPT first extracts highly task-related image features from the visual scene, and then explores the global dependency relationships among all the image regions to generate each image region a global feature. Next, VSPT simulates the human visual working memory to consider all the previous fixations’ influences when predicting each fixation. Experimental findings confirm that our model exhibits adherence to classical visual principles during saccadic decision-making, surpassing the current state-of-the-art performance in free-viewing and task-driven (goal-driven and question-driven) visual scenarios.
Mengyu Qiu, Quan Rong, Dong Liang 0008, Huawei Tu
ISMAR3
2023 Visual Tracking Based on Efficient Dual-Branch Siamese Network
abstract
In this paper, we propose a dual-branch Siamese network for visual object tracking. The proposed network consists of two distinct branches: a shallow network branch and a deep network branch. The shallow network branch focuses on precisely locating the target object and improving the anti-interference ability to similar objects, while the deep network branch focuses on capturing the more abstract semantic features of the object. Additionally, a multi-scale key feature fusion module is embedded into the shallow network, enabling the model to accurately locate the target object. Furthermore, we leverage the attention mechanism to further enhance the robustness of the model. Experimental results on three different public datasets demonstrate that our method outperforms state-of-the-art tracking algorithms.
Dong Liang 0008, Bo Peng 0013
SMC3
2023 Learning rules in spiking neural networks: A survey
Zexiang Yi, Jing Lian 0001, Qidong Liu 0001, Hegui Zhu, Dong Liang 0008, Jizhao Liu
Neurocomputing5
2023 MUS-CDB: Mixed Uncertainty Sampling With Class Distribution Balancing for Active Annotation in Aerial Object Detection
abstract
Recent aerial object detection models rely on a large amount of labeled training data, which requires unaffordable manual labeling costs in large aerial scenes with dense objects. Active learning effectively reduces the data labeling cost by selectively querying the informative and representative unlabelled samples. However, existing active learning methods are mainly with class-balanced settings and image-based querying for generic object detection tasks, which are less applicable to aerial object detection scenarios due to the long-tailed class distribution and dense small objects in aerial scenes. In this paper, we propose a novel active learning method for cost-effective aerial object detection. Specifically, both object-level and image-level informativeness are considered in the object selection to refrain from redundant and myopic querying. Besides, an easy-to-use class-balancing criterion is incorporated to favor the minority objects to alleviate the long-tailed class distribution problem in model training. We further devise a training loss to mine the latent knowledge in the unlabeled image regions. Extensive experiments are conducted on the DOTA-v1.0 and DOTA-v2.0 benchmarks to validate the effectiveness of the proposed method. For the ReDet, KLD, and SASM detectors on the DOTA-v2.0 dataset, the results show that our proposed MUS-CDB method can save nearly 75% of the labeling cost while achieving comparable performance to other active learning methods in terms of mAP. Code is publicly online.
Dong Liang 0008, Jing-Wei Zhang, Ying-Peng Tang, Sheng-Jun Huang
IEEE Trans. Geosci. Remote. Sens.1
2023 Contactless Palmprint Image Recognition Across Smartphones With Self-Paced CycleGAN
abstract
Contactless palmprint recognition, an emerging biometric technology, has attracted increasing attention due to its noninvasive and high practicability characteristics. Although it is naturally suitable for mobile application scenarios, the following two challenges severely limit its recognition performance: 1) the inconsistency in acquisition devices used in training and testing, and 2) many subjects are unable to be imaged on each device, resulting in incomplete data problems. To address these issues, we propose a self-paced CycleGAN with self-attention modules, which simultaneously synthesizes missing data and alleviates the influence of different imaging devices. Specifically, we develop CycleGAN with self-attention modules to generate missing training data by effectively mining the structural correlation among samples while capturing the cross-domain features. Furthermore, a self-paced learning strategy, which is a human cognitive-driven learning mechanism, is used to guide learning the robust cross-domain feature representation and recognition model, by which the relatively easy learning samples are gradually involved in the training process. To verify the effectiveness of the proposed method, we conduct experiments on contactless palmprint datasets collected using different smartphones. The results show that our approach outperforms state-of-the-art methods in classifying contactless palmprint images.
Qi Zhu 0001, Guangnan Xin, Lunke Fei, Dong Liang 0008, Zheng Zhang 0006, Daoqiang Zhang, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.4
2023 Robust RGB-T Tracking via Graph Attention-Based Bilinear Pooling
abstract
RGB-T tracker possesses strong capability of fusing two different yet complementary target observations, thus providing a promising solution to fulfill all-weather tracking in intelligent transportation systems. Existing convolutional neural network (CNN)-based RGB-T tracking methods often consider the multisource-oriented deep feature fusion from global viewpoint, but fail to yield satisfactory performance when the target pair only contains partially useful information. To solve this problem, we propose a four-stream oriented Siamese network (FS-Siamese) for RGB-T tracking. The key innovation of our network structure lies in that we formulate multidomain multilayer feature map fusion as a multiple graph learning problem, based on which we develop a graph attention-based bilinear pooling module to explore the partial feature interaction between the RGB and the thermal targets. This can effectively avoid uninformed image blocks disturbing feature embedding fusion. To enhance the efficiency of the proposed Siamese network structure, we propose to adopt meta-learning to incorporate category information in the updating of bilinear pooling results, which can online enforce the exemplar and current target appearance obtaining similar sematic representation. Extensive experiments on grayscale-thermal object tracking (GTOT) and RGBT234 datasets demonstrate that the proposed method outperforms the state-of-the-art methods for the task of RGB-T tracking.
Bin Kang, Dong Liang 0008, Junxi Mei, Xiaoyang Tan, Dengyin Zhang
IEEE Trans. Neural Networks Learn. Syst.2
2023 Unsupervised inner-point-pairs model for unseen-scene and online moving object detection
Xinyue Zhao, Guangli Wang, Zaixing He, Dong Liang 0008, Shuyou Zhang 0001, Jianrong Tan
Vis. Comput.4
2022 Semantically Contrastive Learning for Low-Light Image Enhancement
abstract
Low-light image enhancement (LLE) remains challenging due to the unfavorable prevailing low-contrast and weak-visibility problems of single RGB images. In this paper, we respond to the intriguing learning-related question -- if leveraging both accessible unpaired over/underexposed images and high-level semantic guidance, can improve the performance of cutting-edge LLE models? Here, we propose an effective semantically contrastive learning paradigm for LLE (namely SCL-LLE). Beyond the existing LLE wisdom, it casts the image enhancement task as multi-task joint learning, where LLE is converted into three constraints of contrastive learning, semantic brightness consistency, and feature preservation for simultaneously ensuring the exposure, texture, and color consistency. SCL-LLE allows the LLE model to learn from unpaired positives (normal-light)/negatives (over/underexposed), and enables it to interact with the scene semantics to regularize the image enhancement network, yet the interaction of high-level semantic knowledge and the low-level signal prior is seldom investigated in previous methods. Training on readily available open data, extensive experiments demonstrate that our method surpasses the state-of-the-arts LLE models over six independent cross-scenes datasets. Moreover, SCL-LLE's potential to benefit the downstream semantic segmentation under extremely dark conditions is discussed. Source Code: https://github.com/LingLIx/SCL-LLE.
Dong Liang 0008, Ling Li 0010, Mingqiang Wei, Wenhan Yang, Huiyu Zhou 0001
AAAI1
2022 I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection
abstract
Can you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to highlight the separator (or say the camouflaged object's boundary) between an image's background and foreground: the reverse attention stream helps erase the camouflaged object's interior to focus on the background, while the normal attention stream recovers the interior and thus pay more attention to the foreground; and both streams are followed by a boundary guider module and combined to strengthen the understanding of boundary. The core design of such separated attention is motivated by the COD procedure of humans: find the subtle difference between the foreground and background to delineate the boundary of a camouflaged object, then the boundary can help further enhance the COD accuracy. We validate on three benchmark datasets that the proposed BSA-Net is very beneficial to detect camouflaged objects with the blurred boundaries and similar colors/patterns with their backgrounds. Extensive results exhibit very clear COD improvements on our BSA-Net over sixteen SOTAs.
Peng Li 0064, Haoran Xie 0001, Xuefeng Yan 0001, Dong Liang 0008, Dapeng Chen, Mingqiang Wei, Harry Qin
AAAI5
2022 MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain Removal
abstract
Rain severely degrades the visibility of scene objects, especially when images are captured through the glass under rainy weather. We observe three intriguing phenomena: 1) rain is a mixture of raindrops, rain streaks and rainy haze; 2) the depth from the camera determines the degree of object visibility, where objects nearby and far away are visually blocked by rain streaks and rainy haze, respectively; and 3) raindrops on the glass randomly affect the object visibility of the whole image space. However, existing solutions and benchmark datasets lack full consideration of the mixture of rain (MOR). In this paper, we originally consider that the overall object visibility is determined by MOR, and enrich the RainCityscapes by considering real-world raindrops to construct the MOR dataset, named RainCityscapes++. To solve the practical rain removal problem arisen from MOR, we formulate a new rain imaging model and propose a multi-branch attention generative adversarial network (MBA-RainGAN). Extensive experiments show clear improvements of our approach over SOTAs on RainCityscapes++.
Yiyang Shen, Yidan Feng, Weiming Wang 0002, Dong Liang 0008, Harry Qin, Haoran Xie 0001, Mingqiang Wei
ICASSP4
2022 More Than Accuracy: An Empirical Study of Consistency Between Performance and Interpretability
Dong Liang 0008, Rong Quan, Songlin Du, Yaping Yan
PRICAI (3)2
2022 GlassNet: Label Decoupling-based Three-stream Neural Network for Robust Image Glass Detection
abstract
Abstract Most of the existing object detection methods generate poor glass detection results, due to the fact that the transparent glass shares the same appearance with arbitrary objects behind it in an image. Different from traditional deep learning‐based wisdoms that simply use the object boundary as an auxiliary supervision, we exploit label decoupling to decompose the original labelled ground‐truth (GT) map into an interior‐diffusion map and a boundary‐diffusion map. The GT map in collaboration with the two newly generated maps breaks the imbalanced distribution of the object boundary, leading to improved glass detection quality. We have three key contributions to solve the transparent glass detection problem: (1) We propose a three‐stream neural network (call GlassNet for short) to fully absorb beneficial features in the three maps. (2) We design a multi‐scale interactive dilation module to explore a wider range of contextual information. (3) We develop an attention‐based boundary‐aware feature Mosaic module to integrate multi‐modal information. Extensive experiments on the benchmark dataset exhibit clear improvements of our method over SOTAs, in terms of both the overall glass detection accuracy and boundary clearness.
Ding Shi, Xuefeng Yan 0001, Dong Liang 0008, Mingqiang Wei, Xin Yang 0011, Yanwen Guo 0001, Haoran Xie 0001
Comput. Graph. Forum4
2022 Inferred box harmonization and aggregation for degraded face detection in crowds
Dong Liang 0008, Qixiang Geng, Huiyu Zhou 0001, Shun'ichi Kaneko
Multim. Tools Appl.1
2022 Anchor Retouching via Model Interaction for Robust Object Detection in Aerial Images
abstract
Object detection has made tremendous strides in computer vision. Small object detection with appearance degradation is a prominent challenge, especially for aerial observations. To collect sufficient positive/negative samples for heuristic training, most object detectors preset region anchors in order to calculate intersection-over-union (IoU) against the ground-truth data. In this case, small objects are frequently abandoned or mislabeled. In this article, we present an effective dynamic enhancement anchor network (DEA-Net) to construct a novel training sample generator. Different from the other state-of-the-art (SOTA) techniques, the proposed network leverages a sample discriminator to realize interactive sample screening between an anchor-based unit and an anchor-free unit to generate eligible samples. Besides, multi-task joint training with a conservative anchor-based inference scheme enhances the performance of the proposed model while reducing computational complexity. The proposed scheme supports both oriented and horizontal object detection tasks. Extensive experiments on two challenging aerial benchmarks (i.e., Dataset of Object deTection in Aerial images (DOTA) and HRSC2016) indicate that our method achieves SOTA performance in accuracy with moderate inference speed and computational overhead for training. On DOTA, our DEA-Net which integrated with the baseline of RoI-transformer surpasses the advanced method by 0.40% mean-average-precision (mAP) for oriented object detection with a weaker backbone network (ResNet-101 vs. ResNet-152) and 3.08% mAP for horizontal object detection with the same backbone. Besides, our DEA-Net which integrated with the baseline of ReDet achieves the SOTA performance by 80.37%. On HRSC2016, it surpasses the previous best model by 1.1% using only three horizontal anchors. The source code and the training set are made publicly available athttps://github.com/QxGeng/DEA-Net.
Dong Liang 0008, Qixiang Geng, Zongqi Wei, Dmitry A. Vorontsov, Ekaterina L. Kim, Mingqiang Wei, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Dense Face Detection via High-level Context Mining
abstract
The appearance degradation caused by low resolution is the core problem of small face detection. Therefore, a natural approach is to assemble information from the context. This paper focuses on how to use high-level contextual information to improve the abilities of anchor-based detectors to detect dense and degenerate faces. We tap the spatial contextual information on the overall view based on the density map, and propose the prior of face co-occurrence for inferred bounding-boxes coordination. We also propose score-size-specific non-maximum suppression to replace the traditional non-maximum suppression at the end of anchor-based detectors. According to the inferred face boxes' quantity, score and size, the proposed synthetical solution reduces false positives and increases true positives. Our method does not require additional training, which is model-independent and can be embedded into existing face detectors. We also propose a dataset - Crowd Face for face detection, which is full of challenges. We expect to supply enough samples to highlight the difficulties of detecting dense and degenerate faces. We embed our proposed methods into state-of-the-art face detectors on massively benchmarked face datasets. Compared with the prior art on the WIDER FACE hard set, our method increase an Average Precision of 0.1 %-1.3%. On Crowd Face, it increases an Average Precision of 1 % – 6%. Dataset is available on: https://github.com/QxGeng/Crowd-Face.
Qixiang Geng, Dong Liang 0008, Huiyu Zhou 0001, Liyan Zhang 0001, Ningzhong Liu
FG2
2021 Robust Spatial-Temporal Correlation Model for Background Initialization in Severe Scene
abstract
Scene background initialization is an important step as one low-layer method for high-layer applications in computer vision. However, this process is always affected by practical challenges such as illumination changes, back-ground motion, camera jitter, intermittent movement and bad weather outdoors, etc. In this work, we develop a novel method called co-occurrence pixel-block (CPB) model via spatial-temporal correlation for robust back-ground initialization. This work first introduces the CPB method for foreground extraction. And then, background information in spatial-temporal features are utilized to recover an adaptive background for the current frame. Experimental results obtained from the dataset of the challenging benchmark (SBMnet) validate it’s performance under various challenges.
Yuheng Deng, Bo Peng 0013, Dong Liang 0008, Shun'ichi Kaneko
ICASSP4
2021 Nlkd: Using Coarse Annotations For Semantic Segmentation Based on Knowledge Distillation
abstract
Modern supervised learning relies on a large amount of training data, yet there are many noisy annotations in real datasets. For semantic segmentation tasks, pixel-level annotation noise is typically located at the edge of an object, while pixels within objects are fine-annotated. We argue the coarse annotations can provide instructive supervised information to guide model training rather than be discarded. This paper proposes a noise learning framework based on knowledge distillation NLKD, to improve segmentation performance on unclean data. It utilizes a teacher network to guide the student network that constitutes the knowledge distillation process. The teacher and student generate the pseudo-labels and jointly evaluate the quality of annotations to generate weights for each sample. Experiments demonstrate the effectiveness of NLKD, and we observe better performance with boundary-aware teacher networks and evaluation metrics. Furthermore, the proposed approach is model-independent and easy to implement, appropriate for integration with other tasks and models.
Dong Liang 0008, Liyan Zhang 0001, Ningzhong Liu, Mingqiang Wei
ICASSP1
2021 Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction
abstract
Using only one deep model for cross scene video foreground segmentation is still very challenging because existing methods are scene-dependent, which restricts the consistent segmentation. In this paper, we propose a cross scene video foreground segmentation framework to extend the generalization capability of those supervised model depending on scene-specific training. The proposed framework flexibly utilizes three well-trained supervised models as guidance to yield a coarse segmentation mask. The co-occurrence probability-based unsupervised background subtraction model is introduced to achieve scene adaptation in the plug and play style without any fine-tuning and labels. Experimental results on LIMU and CDNet2014 datasets validate our framework outperforms the state-of-the-art supervised/unsupervised approaches that participate in the comparison. Experiments also show the training efficiency-related improvements – when introducing the guidance models, the demand for quantity and quality of training samples to train the unsupervised model is reduced. Codes https://github.com/MeteoorLiu/Venus/tree/MeteoorLiu-SUMC
Dong Liang 0008, Bin Kang, Liyan Zhang 0001, Ningzhong Liu
ICASSP1
2021 Robust Cross-Scene Foreground Segmentation in Surveillance Video
abstract
1Training only one deep model for large-scale cross-scene video foreground segmentation is challenging due to the off-the-shelf deep learning based segmentor relies on scene-specific structural information. This results in deep models that are scene-biased and evaluations that are scene-influenced. In this paper, we integrate dual modalities (foregrounds’ motion and appearance), and then eliminating features without representativeness of foreground through attention-module-guided selective-connection structures. It is in an end-to-end training manner and to achieve scene adaptation in the plug and play style. Experiments indicate the proposed method significantly outperforms the state-of-the-art deep models and background subtraction methods in un-trained scenes – LIMU and LASIESTA. Source Code is available at: https://github.com/WeiZongqi/HOFAM
Dong Liang 0008, Zongqi Wei, Huiyu Zhou 0001
ICME1
2021 MPI: Multi-receptive and parallel integration for salient object detection
abstract
Abstract The semantic representation of deep features is essential for image context understanding, and effective fusion of features with different semantic representations can significantly improve the model's performance on salient object detection. This paper proposes a novel method called multi‐receptive and parallel integration, for salient object detection. Firstly, a multi‐receptive enhancement module is designed to effectively expand the receptive fields of features from different layers and generate features with different receptive fields. Multi‐receptive enhancement module can enhance the semantic representation and improve the model's perception of the image context, which enables the model to locate the salient object accurately. Secondly, in order to reduce the reuse of redundant information in the complex top‐down fusion method and weaken the differences between semantic features, a relatively simple but effective parallel fusion strategy is proposed. It allows multi‐scale features to better interact with each other, thus improving the overall performance of the model. Experimental results on multiple datasets demonstrate that the proposed method outperforms state‐of‐the‐art methods under different evaluation metrics.
Jun Cen, Ningzhong Liu, Dong Liang 0008, Huiyu Zhou 0001
IET Image Process.4
2021 Cross-scene foreground segmentation with supervised and unsupervised model communication
Dong Liang 0008, Bin Kang, Pan Gao 0001, Xiaoyang Tan, Shun'ichi Kaneko
Pattern Recognit.1
2020 Coarse-to-fine Foreground Segmentation based on Co-occurrence Pixel-Block and Spatio-Temporal Attention Model
abstract
Foreground segmentation in dynamic scene is an important task in video surveillance. The unsupervised background subtraction method based on background statistics modeling has difficulties in updating. On the other hand, the supervised foreground segmentation method based on deep learning relies on the large-scale of accurately annotated training data, which limits its cross-scene performance. In this paper, we propose a foreground segmentation method from coarse to fine. First, a across-scenes trained Spatio-Temporal Attention Model (STAM) is used to achieve coarse segmentation, which does not require training on specific scene. Then the coarse segmentation is used as a reference to help Co-occurrence Pixel-Block Model (CPB) complete the fine segmentation, and at the same time help CPB to update its background model. This method is more flexible than those deep-learning-based methods which depends on the specific-scene training, and realizes the accurate online dynamic update of the background model. Experimental results on WallFlower and LIMU validate our method outperforms STAM, CPB and other methods of participating in comparison.
Dong Liang 0008
ICPR1
2020 Robust defect detection in 2D images printed on 3D micro-textured surfaces by multiple paired pixel consistency in orientation codes
abstract
Defect detection is now an active research area for production quality assurance. Traditional visual inspection systems are developed by human beings, which is a time‐consuming, labour‐intensive, and highly error‐prone process, and are therefore unreliable. To overcome these problems, the authors proposed a new method for detecting defects when printing on a 3D micro‐textured surface. They utilise an orientation code as the basis to resist the fluctuations in illumination. Based on the consistency of the pixel pairs, they developed a model called multiple paired pixel consistency to represent the statistical relationship between each pixel pair in defect‐free images. Finally, based on this model, they designed a defect detection method. Even with different defect sizes, illumination conditions, noise intensities, and other characteristics, the performance of the proposed algorithm is extremely stable and highly accurate, and the recall, precision, and F‐measure in most of the results can reach 0.85,0.93, and 0.9, respectively. In addition, the defect detection rate can reach almost 100%. This demonstrates that the authors' approach can achieve state‐of‐the‐art accuracy in real industrial applications.
Dong Liang 0008, Shun'ichi Kaneko, Hirokazu Asano
IET Image Process.2
2020 Grayscale-Thermal Tracking via Inverse Sparse Representation-Based Collaborative Encoding
abstract
Grayscale-thermal tracking has attracted a great deal of attention due to its capability of fusing two different yet complementary target observations. Existing methods often consider extracting the discriminative target information and exploring the target correlation among different images as two separate issues, ignoring their interdependence. This may cause tracking drifts in challenging video pairs. This paper presents a collaborative encoding model called joint correlation and discriminant analysis based inver-sparse representation (JCDA-InvSR) to jointly encode the target candidates in the grayscale and thermal video sequences. In particular, we develop a multi-objective programming to integrate the feature selection and the multi-view correlation analysis into a unified optimization problem in JCDA-InvSR, which can simultaneously highlight the special characters of the grayscale and thermal targets through alternately optimizing two aspects: the target discrimination within a given image and the target correlation across different images. For robust grayscale-thermal tracking, we also incorporate the prior knowledge of target candidate codes into the SVM based target classifier to overcome the overfitting caused by limited training labels. Extensive experiments on GTOT and RGBT234 datasets illustrate the promising performance of our tracking framework.
Bin Kang, Dong Liang 0008, Wan Ding, Huiyu Zhou 0001, Wei-Ping Zhu 0001
IEEE Trans. Image Process.2
2019 Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection
abstract
Face detection is an ultimate component to support various visual facial related tasks. However, detecting faces with extremely low resolution or high occlusion is still an open problem. In this paper, we propose a two-step general approach to refine the performance of modern face detectors according to human's high-level context-aware ability. First, we propose Score-specific Non-Maximum Suppression (SNMS) to preserve overlapped faces. Second, we consider the coexistence prior among faces in the scene, which could raise the sensitivity of face detection in the crowd. When integrating our approach to the existing face detectors, most of them have better results on a challenging benchmark (WIDER FACE) and a newly proposed dataset (Faces in Crowd, FIC) made by us. Codes are available on https://github.com/AIoTP/SNMSandCoexistence.
Tianpeng Wu, Dong Liang 0008, Jiaxing Pan, Bin Kang, Shun'ichi Kaneko, Huiyu Zhou 0001
ICASSP2
2019 Context-Anchors for Hybrid Resolution Face Detection
abstract
Despite the positive trends in the development of face detection, open challenges still exist, such as the detection of degraded faces caused by small-size, defocus blur and occlusion in surveillance video. When utilizing anchor-based methods, the anchors are the basic units of training samples, and their ranges are proportional to the ranges of the original label boxes (ground truth). This paper argues that the selected range of an anchor is crucial for a detection task and proposes a face detection model CAHR (context-anchors for hybrid-resolution model) to balance the image resolution and the spatial context range for the purposes of locating small faces. In the training phase, specific size of spatial context is introduced for each anchor, and an image pyramid is employed for a dual CNNs model. In experiments, the indepth analysis of amplification ratio of the anchor and the detection rate is revealed. The detection rate of the small faces is improved by using the proposed model. It is also validated with a massively face datasets (WIDER FACE), demonstrating its superiority to the original hybrid-resolution model (HR) and some other advanced methods.
Tianpeng Wu, Dong Liang 0008, Jiaxing Pan, Shun'ichi Kaneko
ICIP2
2019 Robust visual tracking via nonlocal regularized multi-view sparse representation
Bin Kang, Wei-Ping Zhu 0001, Dong Liang 0008
Pattern Recognit.3
2019 Foreground detection based on co-occurrence background model with hypothesis on degradation modification in dynamic scenes
Shun'ichi Kaneko, Manabu Hashimoto, Yutaka Satoh, Dong Liang 0008
Signal Process.5
2019 Texture-Distortion-Constrained Joint Source-Channel Coding of Multi-View Video Plus Depth-Based 3D Video
abstract
A novel joint source and channel coding scheme tailored to 3D video is proposed in this paper to minimize the end-to-end view synthesis distortion within a given total bit rate for both texture and depth as well as a maximum tolerable distortion constraint for texture. First, we formulate a joint texture and depth coding mode selection strategy for error-resilient source coding of multi-view video plus depth-based 3D video through using the Lagrange multiplier method. Then, by considering the effect of residual errors after channel coding, we evolve to a more general formulation that jointly optimizes error-resilient source coding and channel coding in an integrated manner for unequal error protection between texture and depth, for which a theoretic solution using a proposed dual-trellis is derived. Finally, we extend the general formulation by including the texture distortion constraint. We show how to optimize the view synthesis quality while simultaneously catering to the texture quality constraint. Experimental results demonstrate the proposed algorithm has much better performance than existing related work.
Pan Gao 0001, Wei Xiang 0001, Dong Liang 0008
IEEE Trans. Circuits Syst. Video Technol.3
2018 A Co-occurrence Background Model with Hypothesis on Degradation Modification for Object Detection in Strong Background Changes
abstract
Object detection has become an indispensable part of video processing and current background models are sensitive to background changes. In this paper, we propose a novel background model using an algorithm called Co-occurrence Pixel-block Pairs (CPB) against background changes, such as illumination changes and background motion. We utilize the co-occurrence “pixel to block” structure to extract the spatial-temporal information of each pixel to build background model, and then employ an efficient evaluation strategy to identify the current state of each pixel, which is named as correlation dependent decision function. Furthermore, we also introduce a Hypothesis on Degradation Modification (HoD) into CPB structure to reinforce the robustness of CPB. Experimental results obtained from the dataset of the PETS 2001, AIST-Indoor, SBMnet and CDW-2012 databases show that our models can detect objects robustly in strong background changes.
Shun'ichi Kaneko, Manabu Hashimoto, Yutaka Satoh, Dong Liang 0008
ICPR5
2017 Robust visual tracking via multi-view discriminant based sparse representation
abstract
In traditional sparse representation based visual tracking, the particles are densely sampled, the appearance of some candidates may be very similar, hence the particle observations can be divided into disjointed groups. Existing methods only exploit the group similarity in a certain feature space. In this paper we propose a multi-view discriminant based multitask sparse representation method to exploit the group similarity in a multi-feature space. The proposed method can discriminate the reliability of observation groups and achieve a proper multi-view fusion by using a multi-view discriminant matrix to project multi-feature observation groups into a common subspace. Experiment results show that our method can achieve a better tracking performance than state-of-the-art tracking methods do.
Bin Kang, Dong Liang 0008, Suofei Zhang
ICIP2
2017 Adaptive local spatial modeling for online change detection under abrupt dynamic background
abstract
Change detection is an important theme in video processing. To provide reliable detection results in challenging scenes, traditional methods introduced sophisticated statistical distributions and handcraft spatial features to build background models. In this paper, we develop an intuitive background model based on simple statistical distribution and adaptive spatial correlation among pixels: For each observed pixel, we select a group of supporting pixels with high correlation, and then employ a single Gaussian to model the intensity deviations of each pixel pair. To compensate camera motion and fast adapt to dynamic pattern that coming afterwards, a randomized multichannel on-line updating mechanism is introduced. This observation is robust to abrupt illumination variation and dynamic background. Experimental results using all the video sequences provided by three challenging benchmarks (CDW-2012, CDW-2014 and SABS) validate it outperforms many state-of-the-art methods under various situations.
Dong Liang 0008, Shun'ichi Kaneko, Bin Kang
ICIP1
2017 Robust multi-feature visual tracking via multi-task kernel-based sparse learning
abstract
Feature selection and fusion is of crucial importance in multi‐feature visual tracking. This study proposes a multi‐task kernel‐based sparse learning method for multi‐feature visual tracking. The proposed sparse learning method can discriminate the reliable and unreliable features for optimal multi‐feature fusion through using a Fisher discrimination criterion‐based multi‐objective model to adaptively train the kernel weights of different features such as pixel intensity, edge and texture. To guarantee a robustness of the sparse representation method, a mixed norm is employed in the sparse leaning method to adaptively select correlated particle observations for multi‐task sparse reconstruction. Experimental results show that the proposed sparse learning method can achieve a better tracking performance than state‐of‐the‐art tracking methods do.
Bin Kang, Wei-Ping Zhu 0001, Dong Liang 0008
IET Image Process.3
2015 A sparse-representation-based robust inspection system for hidden defects classification in casting components
Xinyue Zhao, Zaixing He, Shuyou Zhang 0001, Dong Liang 0008
Neurocomputing4
2015 Co-occurrence probability-based pixel pairs background model for robust object detection in dynamic scenes
Dong Liang 0008, Shun'ichi Kaneko, Manabu Hashimoto, Kenji Iwata, Xinyue Zhao
Pattern Recognit.1
2015 Robust pedestrian detection in thermal infrared imagery using a shape distribution histogram feature and modified sparse representation classification
Xinyue Zhao, Zaixing He, Shuyou Zhang 0001, Dong Liang 0008
Pattern Recognit.4
2013 Co-occurrence-based adaptive background model for robust object detection
abstract
An illumination-invariant background model for detecting objects in dynamic scenes is proposed. It is robust in the cases of sudden illumination fluctuation as well as burst moving background. Unlike previous works, it distinguishes objects from a dynamic background using co-occurrence character between a target pixel and its supporting pixels in the form of multiple pixel pairs. Experiments used several challenging datasets that proved the robust performance of object detection in various environments.
Dong Liang 0008, Shun'ichi Kaneko, Manabu Hashimoto, Kenji Iwata, Xinyue Zhao, Yutaka Satoh
AVSS1