Yun Liang 0003

dblp:83/2265-3 · DBLP profile ↗
← Back
47ranked-venue papers
21as first author
29since 2021 · last 2026
0000-0003-0799-0054ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 14 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 8 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FloorPlanFormer: Multi-Task Transformer Network for Floor Plan Recognition with Outer-to-Inner Feature Refinement
abstract
Floor plan recognition requires accurate segmentation and classification of entrance doors, outer contours (walls and windows) and inner contours (various room types) , despite strong spatial dependencies and large stylistic differences between different datasets. To overcome these challenges, we propose FloorPlanFormer, a multi-task learning network divided into three phases: the first phase introduces a Swin Transformer backbone with a pixel decoder to extract fine-grained pixel-level semantics; the second phase employs prompt encoder and mask decoder, and a novel Global Contextual Attention Module (GCAM) is designed to generate clear, high-quality outer contour masks; the third stage uses mask transformer decoder to recognize targets and designs a Masked Feature Refinement Module (MFRM) to accurately delineate the inner contour by modeling the relationship between the local inner and outer contours. Finally, we constructed FloorPlan8K, a dataset containing 8200 images and 77434 instances, on which our model was trained and evaluated, and the results greatly outperformed the state-of-the-art general segmentation methods and specialized methods.
Yun Liang 0003, Run Zheng, Shuai Xie, Yishen Lin
AAAI1
2026 Adjustable Balanced Min-Max Cut for graph clustering
Yun Liang 0003, Qimin Liang, Feiping Nie 0001
Pattern Recognit.1
2026 Nighttime image dehazing via a physics-aware dynamic neural model with progressive contrastive regularization
Yun Liang 0003, Xinjie Xiao, Zihan Zhou 0007, Lianghui Li, Yuhui Quan
Pattern Recognit.1
2026 Deep Underwater Image Quality Assessment via Progressive Physics-Aware Multi-Prior Collaboration
abstract
Underwater image quality assessment (UIQA) is a critical research area, challenged by underwater environments such as wavelength-dependent light attenuation, scattering, and non-uniform illumination. Existing deep learning-based UIQA methods often address these degradations in isolation, neglecting their complex interplay with human perception and lacking explicit modeling of underwater optical phenomena. To address this, we propose PhysIQ-Net, a novel framework that integrates physics-driven principles with progressive multi-prior interaction modeling through three key innovations: First, introduce dual physics-based decomposition that separates images into Backscatter, Transmission, Reflectance, and Illuminance components to capture distinct degradation mechanisms; Second, propose prior-guided dynamic filtering that adapts convolutional kernels to image-specific content using physical priors; and Third, propose physic-informed Cross-Domain Feature Interaction that enables bidirectional collaboration between color-aware and structure-aware representations to model their perceptual inter-dependencies. Extensive experiments across multiple benchmark datasets demonstrate that PhysIQ-Net significantly outperforms existing methods, with ablation studies validating each component’s contribution, providing a robust solution for UIQA.
Zihan Zhou 0007, Jiaxue Lan, Yun Liang 0003, Jing Li 0026, Yong Xu 0007, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.3
2025 ACIL-SED: An Acoustic Clustering and Imbalance Learning Method for Sound Event Detection
abstract
Sound Event Detection (SED) aims to identify and locate specific sound events within audio streams. Despite recent advancements, current methodologies exhibit two fundamental limitations: (1) Insufficient modeling of discriminative temporal-spectral characteristics induces compromised differentiation for acoustically similar events. (2) Failing to address both inter-class and intra-class imbalance problems induces biased classifications toward dominant classes and inactive frames. To tackle these challenges, we propose Acoustic Clustering and Imbalance Learning-based Sound Event Detection (ACIL-SED), which consists of Cluster-Specialized Convolutional Recurrent Neural Network (CS-CRNN) and Dual-Objective Adaptive Balance Loss (DOABLoss). The CS-CRNN employs acoustic clustering-guided specialized sub-models, where each cluster’s sub-model focuses on discriminative feature learning in specific temporal-spectral characteristics, thereby resolving the compromised feature representation inherent in a shared single-model architecture. The DOABLoss adaptively combines mean squared error (MSE) and weighted binary cross-entropy (WBCE) to simultaneously address the issues of inter-class duration imbalance and intra-class active-inactive frame imbalance. Experimental results show that ACIL-SED outperforms state-of-the-art methods in complex acoustic environments.
Cankun Zhong, Yun Liang 0003, Tang Luo, Wing W. Y. Ng
SMC3
2025 A bias extraction and penalty method for robust visual question answering
Bingqian Huang, Zichao Pang, Yun Liang 0003
Appl. Intell.4
2025 Clustering with dynamic bipartite graph learning
Yun Liang 0003, Qimin Liang, Cankun Zhong, Feiping Nie 0001
Neurocomputing1
2025 Parameter-efficient target-specialized audio spectrogram transformers via selective cross-domain knowledge distillation
Yun Liang 0003, Tang Luo, Cankun Zhong
Knowl. Based Syst.1
2025 Enhancing CNN-Based Blind Image Quality Assessment via Deep Cross-Layer Pattern Encoding
abstract
Evaluating image quality without reference images, known as blind image quality assessment (BIQA), is crucial for image communication. Recently, convolutional neural networks (CNNs) have emerged as a prominent BIQA approach due to their feature learning power. Usually, both high-level semantic information and low-level details significantly impact perceived visual quality. However, most existing CNN-based methods focus on high-level semantic information via aggregating features on top of the last convolutional layer into a global descriptor, neglecting the importance of shallow, low-level cues. To address this limitation, this paper proposes a novel approach that exploits local encoding and histogram-based pyramid pooling on crosslayer features produced by a CNN, achieving a joint local and global analysis. Specifically, we introduce a cross-layer pattern encoding model that characterizes features generated along convolutional layers via a soft histogram of local 3D binary patterns. This leads to a highly informative yet compact descriptor for score regression. By building this module into a ResNet backbone, we present an effective BIQA model demonstrating state-ofthe-art performance in extensive experiments on synthetic and authentic datasets.
Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Yun Liang 0003, Jing Li 0026, Patrick Le Callet
IEEE Trans. Multim.4
2024 AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
abstract
Recently, Transformers have been introduced into the field of acoustics recognition. They are pre-trained on large-scale datasets using methods such as supervised learning and semi-supervised learning, demonstrating robust generality——It fine-tunes easily to down-stream tasks and shows more robust performance. However, the predominant fine-tuning method currently used is still full fine-tuning, which involves updating all parameters during training. This not only incurs significant memory usage and time costs but also compromises the model’s generality. Other fine-tuning methods either struggle to address this issue or fail to achieve matching performance. Therefore, we conducted a comprehensive analysis of existing fine-tuning methods and proposed an efficient fine-tuning approach based on Adapter tuning, namely AAT. The core idea is to freeze the audio Transformer model and insert extra learn-able Adapters, efficiently acquiring downstream task knowledge without compromising the model’s original generality. Extensive experiments have shown that our method achieves performance comparable to or even superior to full fine-tuning while optimizing only 7.118% of the parameters. It also demonstrates superiority over other fine-tuning methods.
Yun Liang 0003, Shaojian Qiu
ICASSP1
2024 Hierarchical Temporal Attention and Competent Teacher Network for Sound Event Detection
abstract
Sound event detection identifies specific auditory signal occurrences to recognize the sound event class and its temporal localization. While the Convolutional Recurrent Neural Network with a mean-teacher framework shows impressive SED performance, its effectiveness is hindered by its small receptive field, leading to inadequate consideration of global temporal information and imprecise event boundary localization. Moreover, existing detectors overlook the intricate interplay between temporal and frequency information, compromising detection accuracy. Simultaneously, there is an oversight in interactions between student and teacher models, leading to the teacher conveying inaccurate knowledge to the student. To solve these challenges, this paper proposes a novel robust detector named HTA-CTD, incorporating the Hierarchical Temporal Attention (HTA) and Competent Teacher Network (CTN). HTA introduces an adaptive temporal-frequency feature extraction method, while CTN minimizes reliance on strong labels. Experiments on challenging benchmarks show that our HTA-CTD outperforms the state-of-the-art detector and achieves leading performance.
Yun Liang 0003, Shitong Weng, Shenlong Zheng
ICME2
2024 Memory Matching is Not Enough: Jointly Improving Memory Matching and Decoding for Video Object Segmentation
Jintu Zheng, Yun Liang 0003, Wanchao Su
ICPR (22)2
2024 Boosting Imperceptibility of Adversarial Attacks for Environmental Sound Classification
abstract
As artificial intelligence (AI) continues to advance, AI-based audio systems are becoming increasingly vulnerable to adversarial attacks. However, most current studies overlook the scenes of environmental sounds and the imperceptibility of attack. In response to these, we propose a novel frequency-weighted perturbation algorithm for environmental sounds called the Frequency Psychological Attack Algorithm (FPAA). This innovative algorithm incorporates auditory thresholds with psychoacoustic principles during the perturbation generation process to create highly imperceptible adversarial examples. Extensive experiments conducted on two public datasets using multiple models demonstrate that our FPAA algorithm can produce adversarial audio examples that are not only imperceptible to the human ear but also maintain high offensive capability against AI-based audio systems.
Shaojian Qiu, Xiaokang You, Wei Rong, Lifeng Huang, Yun Liang 0003
ICTAI5
2024 Revitalizing Real Image Deraining via a Generic Paradigm towards Multiple Rainy Patterns
Xin Li 0175, Fan Zhou 0001, Yun Liang 0003, Zhuo Su 0001
IJCAI4
2024 Video object segmentation via couple streams and feature memory
abstract
Abstract In recent years, most video segmentation methods use deep CNN to process the input image, but they did not fully mine the rich intermediate predictions in spatio‐temporal space. And, the segmentation challenges such as occlusion, severe deformation and illumination have not been well solved so far. To alleviate these problems, this paper focuses on constructing multi module network structures that represent multi semantics and proposes a video object segmentation network via coupled‐stream architecture with feature memory mechanism. This network first extracts high‐level semantic features, edge features, long‐term and short‐term stable depth features of the target, and then decode them into the segmentation mask of target. In addition, negative skeleton inhibition and frame interpolation are used to prevent the interference of similar objects and motion blur, respectively. The method has a low GPU memory usage, regardless of the number of object in video. And performs 86.5%and 62.4% in J&F measure on DAVIS 2016 and DAVIS 2017 validation set, without fine‐tuning and online training.
Yun Liang 0003, Xinjie Xiao, Shaojian Qiu, Zhuo Su 0001
IET Image Process.1
2024 An Effective Optimization Method for Fuzzy $k$k-Means With Entropy Regularization
abstract
Fuzzy$k$-Means with Entropy Regularization method (ERFKM) is an extension to Fuzzy$k$-Means (FKM) by introducing a maximum entropy term to FKM, whose purpose is trading off fuzziness and compactness. However, ERFKM often converges to a poor local minimum, which affects its performance. In this paper, we propose an effective optimization method to solve this problem, called IRW-ERFKM. First a new equivalent problem for ERFKM is proposed; then we solve it through Iteratively Re-Weighted (IRW) method. Since IRW-ERFKM optimizes the problem with$k\times 1$instead of$d\times k$intermediate variables, the space complexity of IRW-ERFKM is greatly reduced. Extensive experiments on clustering performance and objective function value show IRW-ERFKM can get a better local minimum than ERFKM with fewer iterations. Through time complexity analysis, it verifies IRW-ERFKM and ERFKM have the same linear time complexity. Moreover, IRW-ERFKM has advantages on evaluation metrics compared with other methods. What's more, there are two interesting findings. One is when we use IRW method to solve the equivalent problem of ERFKM with one factor$\mathbf{U}$, it is equivalent to ERFKM. The other is when the inner loop of IRW-ERFKM is executed only once, IRW-ERFKM and ERFKM are equivalent in this case.
Yun Liang 0003, Qiong Huang 0001, Haoming Chen, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.1
2024 Semi-supervised Video Object Segmentation Via an Edge Attention Gated Graph Convolutional Network
abstract
Video object segmentation (VOS) exhibits heavy occlusions, large deformation, and severe motion blur. While many remarkable convolutional neural networks are devoted to the VOS task, they often mis-identify background noise as the target or output coarse object boundaries, due to the failure of mining detail information and high-order correlations of pixels within the whole video. In this work, we propose an edge attention gated graph convolutional network (GCN) for VOS. The seed point initialization and graph construction stages construct a spatio-temporal graph of the video by exploring the spatial intra-frame correlation and the temporal inter-frame correlation of superpixels. The node classification stage identifies foreground superpixels by using an edge attention gated GCN which mines higher-order correlations between superpixels and propagates features among different nodes. The segmentation optimization stage optimizes the classification of foreground superpixels and reduces segmentation errors by using a global appearance model which captures the long-term stable feature of objects. In summary, the key contribution of our framework is twofold: (a) the spatio-temporal graph representation can propagate the seed points of the first frame to subsequent frames and facilitate our framework for the semi-supervised VOS task; and (b) the edge attention gated GCN can learn the importance of each node with respect to both the neighboring nodes and the whole task with a small number of layers. Experiments on Davis 2016 and Davis 2017 datasets show that our framework achieves the excellent performance with only small training samples (45 video sequences).
Yong Zhang 0029, Shaofan Wang 0001, Yun Liang 0003
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Code Multiview Hypergraph Representation Learning for Software Defect Prediction
abstract
Software defect prediction technology aids the reliability assurance team in identifying defect-prone code and assists the team in reasonably allocating limited testing resources. Recently, researchers assumed that the topological associations among code fragments could be harnessed to construct defect prediction models. Nevertheless, existing graph-based methods only concentrate on features of single-view association, which fail to fully capture the rich information hidden in the code. In addition, software defects may involve multiple code fragments simultaneously, but traditional binary graph structures are insufficient for representing these multivariate associations. To address these two challenges, this article proposes a multiview hypergraph representation learning approach (MVHR-DP) to amplify the potency of code features in defect prediction. MVHR-DP initiates by creating hypergraph structures for each code view, which are then amalgamated into a comprehensive fusion hypergraph. Following this, a hypergraph neural network is established to extract code features from multiple views and intricate associations, thereby enhancing the comprehensiveness of representation in the modeling data. Empirical study shows that the prediction model utilizing features generated by MVHR-DP exhibits superior area under the curve (AUC), F-measure, and matthews correlation coefficient (MCC) results compared to baseline approaches across within-project, cross-version, and cross-project prediction tasks.
Shaojian Qiu, Mengyang Huang, Yun Liang 0003, Chaoda Peng, Yuan Yuan 0004
IEEE Trans. Reliab.3
2023 Global Dilated Attention and Target Focusing Network for Robust Tracking
abstract
Self Attention has shown the excellent performance in tracking due to its global modeling capability. However, it brings two challenges: First, its global receptive field has less attention on local structure and inter-channel associations, which limits the semantics to distinguish objects and backgrounds; Second, its feature fusion with linear process cannot avoid the interference of non-target semantic objects. To solve the above issues, this paper proposes a robust tracking method named GdaTFT by defining the Global Dilated Attention (GDA) and Target Focusing Network (TFN). The GDA provides a new global semantics modeling approach to enhance the semantic objects while eliminating the background. It is defined via the local focusing module, dilated attention and channel adaption module. Thus, it promotes semantics by focusing local key information, building long-range dependencies and enhancing the semantics of channels. Subsequently, to distinguish the target and non-target objects both with rich semantics, the TFN is proposed to accurately focus the target region. Different from the present feature fusion, it uses the template as the query to build a point-to-point correlation between the template and search region, and finally achieves part-level augmentation of target feature in the search region. Thus, the TFN efficiently augments the target embedding while weakening the non-target objects. Experiments on challenging benchmarks (LaSOT, TrackingNet, GOT-10k, OTB-100) demonstrate that the GdaTFT outperforms many state-of-the-art trackers and achieves leading performance. Code will be available.
Yun Liang 0003, Qiaoqiao Li, Fumian Long
AAAI1
2023 Dual-domain Feature Learning and Cross Dimension Interaction Attention for Nighttime Image Dehazing
abstract
Nighttime image dehazing is critical for many computer applications. Directly transferring daytime dehazing models to nighttime scenes often introduces haze residual, detail loss and color distortion for the uneven distribution by artificial lights. Therefore, we propose a nighttime dehazing method by defining the Dual-domain Feature Learning Module (DFLM) and the Feature Optimization Module (FOM). Firstly, we construct the DFLM in both frequency and spatial domains to accurately predict the image degradation caused by haze and remove most haze in nighttime hazy images. Secondly, to address the challenges of uneven illumination distribution and color interference of light sources in nighttime, we construct the FOM based on the proposed Cross Dimension Interaction Attention (CDIA), which captures the feature dependencies by crossing different dimensions including the channel-channel, height-channel and width-channel. By precisely representing illumination and color features, the FOM alleviates color distortion in nighttime dehazing. Extensive experiments on several synthetic and real-world datasets demonstrate that our method outperforms most state-of-the-art methods. Code will be available.
Yun Liang 0003, Xinjie Xiao, Lianghui Li
MMAsia1
2023 SASSM: Semantic Awareness and Self-Support Matching for Semi-Supervised Video Object Segmentation
abstract
Matching-based methods have becamed popular in semi-supervised video object segmentation (VOS), by maintaining a memory bank to predict object masks. However, these methods encounter challenges for fast motions and appearance changes, resulting in blurred predictions and missing boundaries. Then we introduce an innovative network that exploits the self-feature of the query frame to improve the masks prediction. We propose a semantic-aware branch (SAB) for precise semantic guidance during readout decoding and an enhanced feature memory matching module with a self-support matching (SSM) mechanism. Ablations demonstrate the strong collaboration between the semantic-aware branch and the self-support matching mechanism. Our approach achieves a favourable performance on popular datasets, demonstrating a acceptable accuracy and speed performance of 86.3 J&F and 26 FPS on DAVIS 2017 validation. Code will be available.
Yun Liang 0003, Ming Junhui, Jintu Zheng
MMAsia1
2023 GTTrack: Gaussian Transformer Tracker for Visual Tracking
abstract
Recently, Transformer based visual object tracking methods have achieved impressive advancements and significantly improved tracking performance. Transformer includes two modules of self-attention and cross-attention for those methods. However, it brings up two problems: first, the self-attention only considers the relative relation between elements when establishing global association, which can not highlight the essential areas of the tracked target. Second, the cross-attention only relies on feature similarity to locate the target, where the interference of similar objects is challenging. In this paper, we propose a new transformer tracking method of GTTrack by defining Gaussian Attention (GA) and Adaptive Focusing Module (AFM). The GA leads into Gaussian prior to generate a semantic template with robust object features, in which Gaussian prior pays more attention to the central region of the tracked target. The AFM calculates the similarity between current frame and the template by combining the appearance features and position features. The position features are defined with an adaptive Gaussian prior according to the target area in the previous frame. The introduction of position features enhances the contrast between the tracked target and the similar objects. Extensive experiments also demonstrate that the GTTrack outperforms many state-of-the-art trackers and achieves leading performance. Code will be available.
Yun Liang 0003, Fumian Long, Qiaoqiao Li, Dong Wang 0041
MMAsia1
2023 Feature-preserving color pencil drawings from photographs
abstract
Color pencil drawing is well-loved due to its rich expressiveness. This paper proposes an approach for generating feature-preserving color pencil drawings from photographs. To mimic the tonal style of color pencil drawings, which are much lighter and have relatively lower saturation than photographs, we devise a lightness enhancement mapping and a saturation reduction mapping. The lightness mapping is a monotonically decreasing derivative function, which not only increases lightness but also preserves input photograph features. Color saturation is usually related to lightness, so we suppress the saturation dependent on lightness to yield a harmonious tone. Finally, two extremum operators are provided to generate a foreground-aware outline map in which the colors of the generated contours and the foreground object are consistent. Comprehensive experiments show that color pencil drawings generated by our method surpass existing methods in tone capture and feature preservation.
Dong Wang 0041, Guiqing Li, Chengying Gao, Shengwu Fu, Yun Liang 0003
Comput. Vis. Media5
2023 MaskDis R-CNN: An instance segmentation algorithm with adversarial network for herd pigs
abstract
Abstract The current instance segmentation method can achieve satisfactory results in common scenarios. However, under the overlap or partial occlusion between targets caused by the complex scenes, accurate segmentation of pigs remains a challenging task. To address the problem, the authors propose an instance segmentation method based on Mask Scoring region‐based convolutional neural networks (R‐CNN) (MS R‐CNN), which creates the adversarial network called MaskDis in the head branch of MS R‐CNN. The MaskDis is trained as a discriminator using a generative adversarial network, and the MS R‐CNN model is used as a generator during model training. The adversarial training enables the generator to learn context information and features at the pixel level, which effectively improves the segmentation quality under pigs’ overlapping or dense occlusions scenes. Experimental conducted on the pig object segmentation dataset show that the proposed approach achieves a precision of 92.03%, a recall of 92.18%, and an F1 score of 0.9210. Compared with the basic MS R‐CNN model, the approach achieved a 2.25% improvement in precision and 1.18% improvement in F1 score. Furthermore, the improved approach outperformed advanced instance segmentation methods such as YOLACT, Swin Transformer, YOLOv5‐seg, and SOLOv2 on COCO evaluation metrics. These experimental results demonstrate the effectiveness of the proposed approach in instance segmentation of pigs in complex scenes, providing technical support for non‐contact pig automatic management.
Shuqin Tu, Qiantao Zeng, Haofeng Liu, Yun Liang 0003, Zhengxin Huang
IET Image Process.4
2023 Boundary-guided part reasoning network for human parsing
Zhuo Su 0001, Huiqiang Guan, Yuntian Lai, Fan Zhou 0001, Yun Liang 0003
Neurocomputing5
2023 Towards real-world haze removal with uncorrelated graph model
Xiaozhe Meng, Fan Zhou 0001, Yun Liang 0003, Zhuo Su 0001
J. Vis. Commun. Image Represent.4
2022 Feature Dense Relevance Network for Single Image Dehazing
abstract
Existing learning-based dehazing methods do not fully use non-local information, which makes the restoration of seriously degraded region very tough. We propose a novel dehazing network by defining the Feature Dense Relevance module (FDR) and the Shallow Feature Mapping module (SFM). The FDR is defined based on multi-head attention to construct the dense relationship between different local features in the whole image. It enables the network to restore the degraded local regions by non-local information in complex scenes. In addition, the raw distant skip-connection easily leads to artifacts while it cannot deal with the shallow features effectively. Therefore, we define the SFM by combining the atmospheric scattering model and the distant skip-connection to effectively deal with the shallow features in different scales. It not only maps the degraded textures into clear textures by distant dependence, but also reduces artifacts and color distortions effectively. We introduce contrastive loss and focal frequency loss in the network to obtain a realitic and clear image. The extensive experiments on several synthetic and real-world datasets demonstrate that our network surpasses most of the state-of-the-art methods.
Yun Liang 0003, Enze Huang, Zhuo Su 0001, Dong Wang 0041
IJCAI1
2021 Multi-features guided robust visual tracking
Yun Liang 0003, Jian Zhang 0026, Mei-hua Wang, Chen Lin 0001
Multim. Tools Appl.1
2021 Robust Visual Tracking Based on Convolutional Sparse Coding
abstract
This paper proposes a new visual tracking method by constructing the robust appearance model of the target with convolutional sparse coding. First, our method uses convolutional sparse coding to divide the interest region of the target into a smooth image and four detail images with different fitting degrees. Second, we compute the initial target region by tracking the smooth image with the kernel correlation filtering. We define an appearance model to describe the details of the target based on the initial target region and the combination of four detail images. Third, we propose a matching method by the overlap rate and Euclidean distance to evaluate candidates and the appearance model to compute the tracking results based on detail images. Finally, the two tracking results are separately computed by the smooth image, and the detail images are combined to produce the final target rectangle. Many experiments on videos from Tracking Benchmark 2015 demonstrate that our method produces much better results than most of the present visual tracking methods.
Yun Liang 0003, Dong Wang 0041, Lei Xiao 0010, Caixing Liu
Wirel. Commun. Mob. Comput.1
2020 Single image rain removal with reusing original input squeeze-and-excitation network
abstract
In this study, the authors propose a novel network architecture to address the problem of removing rain streaks from single images. To strengthen the representational power of the network, they adopt the squeeze‐and‐excitation block in the network. Furthermore, they propose a new network connection called reusing original input (ROI). The ROI connection reuses the original input of the network and can provide more texture details of the background. These details can be useful for the restoration of the image after removing the rain streaks. Batch normalisation is applied to further improve the rain removal performance of the network. Despite the fact that the network is trained on synthetic data, experimental results show that the proposed network has a comparable performance on both synthetic images and real‐world images to the state‐of‐the‐art methods.
Meihua Wang, Lunbao Chen, Yun Liang 0003, Yuexing Hao, Haijun He
IET Image Process.3
2020 On large appearance change in visual tracking
Yun Liang 0003, Meihua Wang, Yanwen Guo 0001, Wei-Shi Zheng 0001
Neural Comput. Appl.1
2019 Drug Target Interaction Prediction using Multi-task Learning and Co-attention
abstract
Various machine learning models have been proposed as cost-effective means to predict Drug-Target Interactions (DTI). Most existing researches treat DTI prediction either as a classification task (i.e. output negative or positive labels to indicate existence of interaction) or as a regression task (i.e. output numerical values as the strength of interaction). However, classifiers are more prone to higher bias and regression models tend to overfit the training data to generate large variance. In this paper, we explore to balance the bias and variance by a multi-task learning framework. We propose an architecture to both predict accurate values of strength of interaction and decide correct boundary between positive and negative interactions. Furthermore, the two tasks are performed on a shared feature representation, which is learnt using a co-attention mechanism. Comprehensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art methods.
Yuyou Weng, Chen Lin 0001, Xiangxiang Zeng, Yun Liang 0003
BIBM4
2019 Generic interactive pixel-level image editing
abstract
Abstract Several image editing methods have been proposed in the past decades, achieving brilliant results. The most sophisticated of them, however, require additional information per‐pixel. For instance, dehazing requires a specific transmittance value per pixel, or depth of field blurring requires depth or disparity values per pixel. This additional per‐pixel value is obtained either through elaborated heuristics or through additional control over the capture hardware, which is very often tailored for the specific editing application. In contrast, however, we propose a generic editing paradigm that can become the base of several different applications. This paradigm generates both the needed per‐pixel values and the resulting edit at interactive rates, with minimal user input that can be iteratively refined. Our key insight for getting per‐pixel values at such speed is to cluster them into superpixels, but, instead of a constant value per superpixel (which yields accuracy problems), we have a mathematical expression for pixel values at each superpixel: in our case, an order two multinomial per superpixel. This leads to a linear least‐squares system, effectively enabling specific per‐pixel values at fast speeds. We illustrate this approach in three applications: depth of field blurring (from depth values), dehazing (from transmittance values) and tone mapping (from brightness and contrast local values), and our approach proves both favorably interactive and accurate in all three. Our technique is also evaluated with a common dataset and compared favorably.
Yun Liang 0003, Yibo Gan, Mingqin Chen, Diego Gutierrez, Adolfo Muñoz 0001
Comput. Graph. Forum1
2019 Robust visual tracking via identifying multi-scale patches
Yun Liang 0003, Ke Li 0005, Jian Zhang 0026, Meihua Wang, Chen Lin 0001
Multim. Tools Appl.1
2018 Drug Target Interaction Prediction with Non-random Missing Labels
Sheng Ni, Chen Lin 0001, Xiangxiang Zeng, Yun Liang 0003
BIBM4
2018 A component-driven distributed framework for real-time video dehazing
Meihua Wang, Jiaming Mai, Yun Liang 0003, Ruichu Cai, Tom Z. J. Fu
Multim. Tools Appl.3
2018 Single image deraining using deep convolutional networks
Meihua Wang, Jiaming Mai, Ruichu Cai, Yun Liang 0003, Hua Wan
Multim. Tools Appl.4
2018 Multi-modal feature fusion for geographic image annotation
Ke Li 0005, Changqing Zou, Shuhui Bu, Yun Liang 0003, Jian Zhang 0026, Minglun Gong
Pattern Recognit.4
2018 A semi-supervised framework for topology preserving performance-driven facial animation
Jian Zhang 0026, Yun Liang 0003
Signal Process.3
2017 Learning 3D faces from 2D images via Stacked Contractive Autoencoder
Jian Zhang 0026, Ke Li 0005, Yun Liang 0003
Neurocomputing3
2017 Objective Quality Prediction of Image Retargeting Algorithms
abstract
Quality assessment of image retargeting results is useful when comparing different methods. However, performing the necessary user studies is a long, cumbersome process. In this paper, we propose a simple yet efficient objective quality assessment method based on five key factors: i) preservation of salient regions; ii) analysis of the influence of artifacts; iii) preservation of the global structure of the image; iv) compliance with well-established aesthetics rules; and v) preservation of symmetry. Experiments on the RetargetMe benchmark, as well as a comprehensive additional user study, demonstrate that our proposed objective quality assessment method outperforms other existing metrics, while correlating better with human judgements. This makes our metric a good predictor of subjective preference.
Yun Liang 0003, Yong-Jin Liu 0001, Diego Gutierrez
IEEE Trans. Vis. Comput. Graph.1
2016 Ensemble-driven support vector clustering: From ensemble learning to automatic parameter estimation
abstract
Support vector clustering (SVC) is a versatile clustering technique that is able to identify clusters of arbitrary shapes by exploiting the kernel trick. However, one hurdle that restricts the application of SVC lies in its sensitivity to the kernel parameter and the trade-off parameter. Although many extensions of SVC have been developed, to the best of our knowledge, there is still no algorithm that is able to effectively estimate the two crucial parameters in SVC without supervision. In this paper, we propose a novel support vector clustering approach termed ensemble-driven support vector clustering (EDSVC), which for the first time tackles the automatic parameter estimation problem for SVC based on ensemble learning, and is capable of producing robust clustering results in a purely unsupervised manner. Experimental results on multiple real-world datasets demonstrate the effectiveness of our approach.
Dong Huang 0001, Chang-Dong Wang 0001, Jian-Huang Lai, Yun Liang 0003, Shan Bian
ICPR4
2015 3D Model retrieval based on skeleton
abstract
We proposed a new method for 3D model retrieval based on skeletons. As we know a sketch is drawn on a two dimensional plane while models are three dimensional. For better comparison we choose the same dimension for them. For simpler user input we use the front view of the skeleton to represent 3D models, in this way 3D models can be represented in 2D form. We get the skeleton of model through Skeleton Extraction algorithm based on mesh simplification and mesh contraction and get the front view of the skeleton by mapping to two dimensional spaces. The model compared with the Model libraries by Feature description and matching to find similar models on shape and topology structure after getting the feature extraction of the model by the gradient histogram matching algorithm. Experiments showed that most users can find the 3D models they want with our system.
Shujin Lin, Yihui Guo, Yun Liang 0003, Yanhua Wu
NAS3
2013 Optimised image retargeting using aesthetic-based cropping and scaling
abstract
Image retargeting is a critical technique in displaying images on devices with different resolutions. This study presents a new image retargeting algorithm based on aesthetic‐based cropping and scaling. A composite measurement is first constructed under the guidelines of composition aesthetics in photographing. An aesthetic‐based cropping is proposed to yield an optimal candidate retargeted image with maximum aesthetic value computed via a constructed composite measurement. The optimal candidate is uniformly scaled to obtain the retargeted image of target size. Some subjective and objective assessments demonstrate that the proposed scheme significantly improves the aesthetics of retargeted images while preserving the important objects. It also achieves better performance in terms of aesthetics than a number of conventional image retargeting approaches.
Yun Liang 0003, Zhuo Su 0001, Chuntao Wang, Dong Wang 0041
IET Image Process.1
2013 Edge-Preserving Texture Suppression Filter Based on Joint Filtering Schemes
abstract
Obtaining a texture-smoothing and edge-preserving filtered output is significant to image decomposition. Although the edge and the texture have salient difference in human vision, automatically distinguishing them is a difficult task, for they have similar intensity difference or gradient response. The state-of-the-art edge-preserving smoothing (EPS) based decomposition approaches are hard to obtain a satisfactory result. We propose a novel edge-preserving texture suppression filter, exploiting the joint bilateral filter as a bridge to achieve the purpose of both properties of texture-smoothing and edge-preserving. We develop the iterative asymmetric sampling and the local linear model to produce the degenerative image to suppress the texture, and apply the edge correction operator to achieve edge-preserving. An efficient accelerating implementation is introduced to improve the performance of filtering response. The experiments demonstrate that our filter produces satisfactory outputs with both properties of texture-smoothing and edge-preserving, while compared with the results of other popular EPS approaches in signal, visual and time analysis. Finally, we extend our filter to a variety of image processing applications.
Zhuo Su 0001, Zhengjie Deng, Yun Liang 0003, Zhen Ji
IEEE Trans. Multim.4
2012 Image Resizing Based on Geometry Preservation with Seam Carving
abstract
When an image or a video is transformed to an aspect ratio deferent from its original size, information lost is inevitable no matter what method is used, thus, how to keep the most attractive contents and minimize the visual distortion during the resizing process is the key issue. To address this problem, this paper proposes an object geometry preservation method based on the seam carving method. We first define a framework that measures the importance of geometry feature in the source material, then a new energy function is presented with object geometry constraint, according to the new energy function, an optimized seam carving method is used to minimize distortion while resizing the source material. The experiment results show that our method is better to transform a variety of source images to a different display size than conventional resizing methods.
Fan Zhou 0001, Ruomei Wang 0001, Yun Liang 0003
TrustCom4
2012 Patchwise scaling method for content-aware image resizing
Yun Liang 0003, Zhuo Su 0001
Signal Process.1