Wenxin Yu 0001

dblp:36/8502-1 · DBLP profile ↗
← Back
101ranked-venue papers
4as first author
76since 2021 · last 2026
0000-0002-6093-5516ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 2 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 1 first-author · 21 since 2021Systems, architecture and hardware · 24 · 1 first-author · 22 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Timing-Driven Detailed Placement with Collaborative Topology Reconstruction
abstract
Placement is a critical step in the physical design, as it largely determines the potential for subsequent optimization. In this work, we propose a timing-driven detailed placement framework: first, a simplified RC-tree model is employed for flip-flop–buffer compensation; then, a gradient-augmented global heuristic algorithm is incorporated; and finally, timing improvement is achieved through local collaborative optimization. A comprehensive evaluation on eight ICCAD 2015 benchmarks demonstrates the effectiveness of our approach. Compared to DREAMPlace4.0-DP, a state-of-the-art timing-driven placer, our framework achieves an average improvement of 19.60% in WNS and 55.74% in TNS, while introducing less disturbance to the global placement. Moreover, it delivers a 0.80% reduction in HPWL and reduces runtime by 20.24%.
Zhengjie Zhao, Wenxin Yu 0001, Mengshi Gong, Youzhi Zheng, Xinmiao Li, Wenyu Liu 0018, Jingwei Lu
DATE2
2025 Uncertainty-Aware Multi-Branch Distillation for Label-Scarce Medical Image Segmentation
abstract
Semi-supervised learning (SSL) has gained attention in medical image segmentation by leveraging abundant unlabeled data to reduce reliance on expert annotations. However, existing methods underutilize the complementary information from diverse augmented views. This information is crucial for handling the inherent variability and subtle pathological changes in medical imaging. To address this issue, we introduce Mixstill, a novel dual-component model comprising Mask-Distillation and Uncertainty Estimation. Mixstill enforces cross-view consistency on multiple data streams under diverse perturbations, promoting robust feature learning that enhances segmentation accuracy under anatomical variability and subtle pathological variations. Building upon this cross-view consistency, we further introduce an uncertainty-aware fusion mechanism that adaptively combines predictions from different views based on their reliability, producing high-quality supervision signals for unlabeled data. Extensive experiments conducted on two public datasets demonstrate the effectiveness of the proposed method. Under scarce labeled data conditions, our approach achieves superior performance compared to state-of-the-art SSL methods.
Liangjie Wang, Ning Jiang 0002, Changzheng Yang, Wenxin Yu 0001
BIBM4
2025 An Effective and Efficient Cross-Link Insertion for Non-Tree Clock Network Synthesis
abstract
Clock skew introduces significant challenge to the overall system performance. Existing non-tree solutions like cross-link insertion often come with limitations, such as the over-consumption of resource and power. In this work, we propose a cross-link insertion algorithm that effectively reduces the clock skew with minimal power overhead, and prioritize delay optimization on the paths with high sensitivity to the skew. The experimental results from the ISPD 2010 benchmarks show a 17% reduction in the mean of clock skew, a 45% decrease in the standard deviation of clock skew, and a 13% lower power consumption versus the advanced non-tree solutions in literature.
Jinghao Ding, Jiazhi Wen, Zhaoqi Fu, Mengshi Gong, Yuanrui Qi, Wenxin Yu 0001, Jinjia Zhou
DATE7
2025 Timing-Driven Global Placement With Hybrid Heuristics and Nadam-Based Net Weighting
abstract
Timing optimization is critical to the entire design flow of the very-large-scale integrated (VLSI) circuit, and Global Placement is pivotal in achieving timing closure within the design flow of very-large-scale integration circuits. However, most global placement algorithms focus on optimizing wirelength rather than timing. Therefore, we propose a novel timing-driven global placement algorithm to address this gap. This paper proposes a timing-driven global placement algorithm utilizing a Nadam-based net-weighting strategy. Additionally, we employ a hybrid heuristic approach for adaptive dynamic adjustment of net weights. The experimental results on the ICCAD 2015 contest benchmarks show that compared to the RePlAce, our algorithm significantly improves WNS and TNS by 40.7% and 56.5%, respectively.
Linhao Lu, Wenxin Yu 0001, Hongwei Tian, Chengjin Li, Xinmiao Li, Zhaoqi Fu, Zhengjie Zhao, Jingwei Lu
DATE2
2025 An Effective Macro Placement Framework with Reinforcement Learning and Monte Carlo Tree Search
Jinghao Ding, Wenxin Yu 0001, Yuanrui Qi, Zhaoqi Fu, Mengshi Gong, I-Chyn Wey, Jinjia Zhou
ACM Great Lakes Symposium on VLSI2
2025 Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed Scenarios
abstract
Knowledge Distillation (KD) transfers knowledge from a large pre-trained teacher network to a compact and efficient student network, making it suitable for deployment on resource-limited media terminals. However, traditional KD methods require balanced data to ensure robust training, which is often unavailable in practical applications. In such scenarios, a few head categories occupy a substantial proportion of examples. This imbalance biases the trained teacher network towards the head categories, resulting in severe performance degradation on the less represented tail categories for both the teacher and student networks. In this paper, we propose a novel framework called Knowledge Rectification Distillation (KRDistill) to address the imbalanced knowledge inherited in the teacher network through the incorporation of the balanced category priors. Furthermore, we rectify the biased predictions produced by the teacher network, particularly focusing on the tail categories. Consequently, the teacher network can provide balanced and accurate knowledge to train a reliable student network. Intensive experiments conducted on various long-tailed datasets demonstrate that our KRDistill can effectively train reliable student networks in realistic scenarios of data imbalance.
Xinlei Huang, Jialiang Tang, Xubin Zheng, Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002
ICASSP5
2025 Symmetry and Fusion Data Augmentation for Semi-Supervised Medical Segmentation
abstract
In semi-supervised medical image segmentation, appropriately merging labeled and unlabeled data before network training instead of using them separately can effectively reduce knowledge loss, mitigate distribution discrepancies and promote efficient knowledge transfer to unlabeled data. However, existing methods tend to focus on data fusion, overlooking the importance of self-transformations. Therefore, we propose a comprehensive data augmentation strategy that combines self-transformations with cross-sample fusion: Symmetry and Fusion Data Augmentation (SF-DA), implemented within the Mean Teacher framework. Our method has two branches: Self-Symmetric Flipping (SSF), enhancing the model’s feature understanding, and cross-sample stitching (CSS), promoting common semantic learning between labeled and unlabeled data. Together, they promote more comprehensive knowledge transfer. Extensive experiments on ACDC and PROMISE12 datasets demonstrate the effectiveness and superiority of SF-DA. Across different labeled data scenarios, SF-DA consistently outperforms the second-best method in all evaluation metrics. Code is available at https://github.com/ZYS-four/SF-DA.git.
Yishan Zhang, Wenxin Yu 0001, Jun Gong 0001
ICASSP2
2025 Timing-Driven Global Placement with Entropy-Mobility Guided Pin-to-Pin Weighting
abstract
Timing-driven global placement is a critical phase in modern very-large-scale integration design, where minimizing total negative slack and worst negative slack is essential for achieving reliable timing closure. While prior works leverage pin-to-pin attraction and path-level timing analysis for optimization, their use of static weighting schemes limits adaptability to diverse critical path characteristics. In this paper, we present an enhanced timing-driven placement framework that introduces a novel entropy and node-aware pin-to-pin weighting strategy. For each pin pair, the weight is computed using its path entropy and the average number of logic nodes across associated critical paths, reflecting both slack concentration and path complexity. Furthermore, a dynamic balancing factor is incorporated to adaptively modulate the contribution of these two components during the optimization process. Experimental results on the ICCAD 2015 benchmark suite demonstrate that our approach significantly outperforms DREAMPlace 4.0, achieving an average improvement of$\mathbf{5 9. 3 9 \%}$in TNS and$\mathbf{2 7. 1 5 \%}$in WNS, along with noticeable reductions in half-perimeter wirelength.
Youzhi Zheng, Zhengjie Zhao, Linhao Lu, Wenxin Yu 0001, Jingwei Lu
ICCD5
2025 Beyond Feature Mapping GAP: Integrating Real HDRTV Priors for Superior SDRTV-to-HDRTV Conversion
abstract
The rise of HDR-WCG display devices has highlighted the need to convert SDRTV to HDRTV, as most video sources are still in SDR. Existing methods primarily focus on designing neural networks to learn a single-style mapping from SDRTV to HDRTV. However, the limited information in SDRTV and the diversity of styles in real-world conversions render this process an ill-posed problem, thereby constraining the performance and generalization of these methods. Inspired by generative approaches, we propose a novel method for SDRTV to HDRTV conversion guided by real HDRTV priors. Despite the limited information in SDRTV, introducing real HDRTV as reference priors significantly constrains the solution space of the originally high-dimensional ill-posed problem. This shift transforms the task from solving an unreferenced prediction problem to making a referenced selection, thereby markedly enhancing the accuracy and reliability of the conversion process. Specifically, our approach comprises two stages: the first stage employs a Vector Quantized Generative Adversarial Network to capture HDRTV priors, while the second stage matches these priors to the input SDRTV content to recover realistic HDRTV outputs. We evaluate our method on public datasets, demonstrating its effectiveness with significant improvements in both objective and subjective metrics across real and synthetic datasets.
Gang He 0002, Kepeng Xu, Li Xu 0008, Wenxin Yu 0001, Xianyun Wu
IJCAI4
2025 Unleashing the Potential of Transformer Flow for Photorealistic Face Restoration
abstract
Face restoration is a challenging task due to the need to remove artifacts and restore details. Traditional methods usually use generative model prior to achieve face restoration, but the restored results are still insufficient in terms of realism and details. In this paper, we introduce OmniFace, a novel face restoration framework that leverages Transformer-based diffusion flow. By exploiting the scaling property of Transformer, OmniFace achieves high-resolution restoration with exceptional realism and detail. The framework integrates three key components: (1) a Transformer-driven vector estimation network, (2) a representation aligned ControlNet, and (3) an adaptive training strategy for face restoration. The inherent scaling law of Transformer architectures enables the restoration of high-quality faces at high resolution. The controlnet combined with pre-trained diffusion representation can be easily trained. The adaptive training strategy provides a vector field that is more suitable for face restoration. Comprehensive experiments demonstrate that OmniFace outperforms existing techniques in terms of restoration quality across multiple benchmark datasets, especially in restoring photographic-level texture details in high-resolution scenes.
Kepeng Xu, Li Xu 0008, Gang He 0002, Wei Chen 0062, Xianyun Wu, Wenxin Yu 0001
IJCAI6
2025 EMIFS: Efficient Multi-scale Information Fusion Self-supervision for Medical Image Segmentation
abstract
Medical image segmentation plays an important role in clinical decision making and auxiliary diagnosis. Today, however, it still faces three major challenges. 1. In the task of medical image segmentation, due to the different types of lesions and the large difference in the size of the lesion area, the segmentation accuracy is seriously reduced. 2. In order to pursue the segmentation performance, the model is difficult to be applied to the actual medical environment due to the excessive parameters. 3. Relying too much on manually labeled images to assist training. In order to meet these challenges, we propose a lightweight segmentation network, which is dedicated to extracting local and global information and fusing multi-level and multi-source features to maximize the segmentation accuracy for different shape lesions, especially for the case of fuzzy boundary and small segmentation target. The method of generating intermediate mask self-monitoring is used to generate additional labeled images to assist training. Finally, by using efficient down sampling and up sampling operations, the parameter quantity is only 1.37M while effectively extracting information. On the BUSI and ISIC2018 datasets, mIoU and DSC scores reached 75.57%, 83.57% and 83.85%, 90.38% respectively, indicating that we have reached the best balance between parameters and performance. The code is available at https://github.com/Jay217219/EMIFS.
Luyao Ren, Wenxin Yu 0001
ACM Multimedia2
2025 Bi-affine Semantic Fusion Generative Adversarial Networks for Text-to-Image Synthesis
abstract
The goal of text-to-image synthesis is to generate realistic images semantically consistent with text. While diffusion models excel in performance, they demand massive data and parameters, causing high power consumption and low deployment efficiency. In contrast, GANs, with compact structures and fast inference, suit low-power designs but suffer from conditional batch normalization issues—inducing feature confusion, disrupting text-image consistency, and limiting semantic fusion. To address this, we propose Bi-Affine Semantic Fusion GAN (BA-GAN). It deepens image-text fusion via channel-spatial affine transformations, enhancing realism and semantic capture. Contrastive loss between text, synthetic, and real images further boosts consistency. BA-GAN also maintains lightweight parameters. Experiments on CUB and COCO validate its pretty improvements in realism, semantic alignment, and efficiency.
Wenxin Yu 0001, Xin Cheng 0004
MMAsia2
2024 A P&R Co- Optimization Engine for Reducing Congestion
abstract
Placement and routing (P&R) are two crucial stages in the physical design process to optimize different objectives. For instance, placement often focuses on optimizing the half-perimeter wirelength (HPWL) and estimated congestion while routing attempts to minimize the wirelength and the number of overflows. The misalignment of objectives inevitably leads to a significant decline in solution quality. Therefore, this paper is an efficient Formula R co-optimization engine that bridges the gap between placement and routing disparities. To progressively alleviate routing congestion issues, we perform cell movements on cells causing overflows and then re-run the routing process. Comparing our experimental results with CUGR [5], our method reduced 17.1 in overflow, with slight decreases in wirelength and vias.
Dongliang Xia, Wenxin Yu 0001, Zhaoqi Fu, Zejun Gan, Chengjin Li
ACM Great Lakes Symposium on VLSI2
2024 Similarity Knowledge Distillation with Calibrated Mask
abstract
In this paper, we propose a novel and efficient method for knowledge distillation, which is structurally simple and requires negligible computation overhead. Our method includes three modules. The first module is the calibrated mask, which avoids the teacher model’s incorrect representation to disturb the student model’s training; the second module and the third module improve the performance of the student model by the similarity of the sample and the process, respectively. The student model attains better performance in qualitative and quantitative evaluation through the judicious amalgamation of these three modules. Our method is experimented with through rigorous validation of canonical datasets, including CIFAR-100 and TinyImageNet. The experimental corroboration conclusively attests to the better performance of our method, soaring above the extant most state-of-the-art on both subjective and objective dimensions.
Qi Wang 0089, Wenxin Yu 0001, Lu Che, Jun Gong 0001
ICASSP2
2024 Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
Kepeng Xu, Li Xu 0008, Gang He 0002, Wenxin Yu 0001, Yunsong Li 0001
IJCAI4
2024 Decoupled Multi-teacher Knowledge Distillation based on Entropy
abstract
Multi-teacher knowledge distillation (MKD) aims to leverage the valuable and diverse knowledge presented by multiple teacher networks to improve the performance of the student network. Existing approaches typically rely on simple methods such as averaging the prediction logits or using sub-optimal weighting strategies to combine knowledge from multiple teachers. However, employing these techniques cannot fully reflect the importance of teachers and may even mislead student’s learning. To address these issues, we propose a novel Decoupled Multi-teacher Knowledge Distillation based on Entropy (DE-MKD). DE-MKD decomposes the vanilla KD loss and assigns weights to each teacher to reflect its importance based on the entropy of their predictions. Furthermore, we extend the proposed approach to distill the intermediate features from teachers to further improve the performance of the student network. Extensive experiments conducted on the publicly available CIFAR-100 image classification dataset demonstrate the effectiveness and flexibility of our proposed approach.
Xin Cheng 0004, Jialiang Tang, Wenxin Yu 0001, Ning Jiang 0002, Jinjia Zhou
ISCAS4
2024 An Attention Network With Self-Supervised Learning for Rheumatoid Arthritis Scoring
abstract
Rheumatoid arthritis (RA) is a chronic disease causing joint pain and disability. Early treatment can prevent irreversible bone damage. The Sharp/van der Heijde (SvH) method, a common clinical standard for assessing RA progression, is manual and subjective, leading to inconsistent measurements. To overcome this, we propose a deep learning model based on the SvH scoring criteria to predict SvH scores of RA patients’ hands. Our efficient attention network, ECBANet, merges supervised and self-supervised learning to generate attentional weights without losing feature dimensionality. It effectively captures the local features of the RA hand and enhances the model’s focus on crucial cues like joint position and gap. Extensive experiments on the only public dataset of hand X-ray images of RA patients labeled with SvH scores show that our model surpasses the current optimal model for this dataset, reducing MAE by 2% and RMSE by 7.4%.
Deyu Ling, Wenxin Yu 0001, Jinmei Zou
ISCAS2
2024 Axial Attention Transformer for Fast High-quality Image Style Transfer
abstract
Image style transfer aims to blend the content of one image with the style of another. Due to the limitation of traditional Convolutional Neural Network (CNN) methods in capturing global information, researchers have turned to vision transformers for image style transfer tasks, aiming to achieve a broader receptive field. However, the large computations of vision transformers lead to longer inference time compared to many conventional CNN-based methods, thereby constraining their practical application. To tackle this challenge, we propose an axial attention transformer encoder named AATE and design a fast vision transformer image style transfer model. Our model encompasses two AATEs that individually process content and style images, along with a transformer decoder to amalgamate style and content. In addition, we use a downsample module to preprocess the network inputs and an upsample module to refine the output. As a result, we improve the inference speed about 5 times faster than current state-of-the-art (SOTA) transformerbased methods while achieving excellent image quality. Qualitative and quantitative experiments show competitive results compared with advanced methods.
Wenxin Yu 0001, Qi Wang 0089, Lu Che
ISCAS2
2024 Track Assignment Using Gradient Indication and Simulated Annealing
abstract
Track assignment has become a crucial step within the physical design flow. In this work, we proposed a novel track assignment approach using gradients estimation and simulated annealing. Specifically, we formulate the overlap cost function and derive the gradients to estimate the quality of candidate positions for iroutes movement. A simulated annealer is devised to consume the gradients indication and perturb the track assignment solution from the global perspective. Experimental results show 5.15% lower overlap cost in average of all 10 benchmarks from DAC 2012 Routability-Driven Placement Contest, compared to the state-of-the-art approach NTA [1].
Yuanrui Qi, Zejun Gan, Jinghao Ding, Zhaoqi Fu, Mengshi Gong, Wenxin Yu 0001
ISCAS6
2024 QS-NeRV: Real-Time Quality-Scalable Decoding with Neural Representation for Videos
abstract
In this paper, we propose a neural representation for videos that enables real-time quality-scalable decoding, called QS-NeRV. QS-NeRV comprises a Self-Learning Distribution Mapping Network (SDMN) and Extensible Enhancement Networks (EENs). Firstly, SDMN functions as the base layer (BL) for scalable video coding, focusing on encoding videos of lower quality. Within SDMN, we employ a methodology that minimizes the bitstream overhead to achieve efficient information exchange between the encoder and decoder instead of direct transmission. Specifically, we utilize an invertible network to map the multi-scale information obtained from the encoder to a specific distribution. Subsequently, during the decoding process, this information is recovered from a randomly sampled latent variable to assist the decoder in achieving improved reconstruction performance. Secondly, EENs serve as the enhancement layers (ELs) and are trained in an overfitting manner to obtain robust restoration capability. By integrating the fixed BL bitstream with the parameters of EEN as an extension pack, the decoder can produce higher-quality enhanced videos. Furthermore, the scalability of the method allows for adjusting the number of combined packs to accommodate diverse quality requirements. Experimental results demonstrate our proposed QS-NeRV outperforms the state-of-the-art real-time decoding INR-based methods on various datasets for video compression and interpolation tasks.
Chang Wu 0001, Guancheng Quan, Gang He 0002, Yunsong Li 0001, Wenxin Yu 0001, Xianmeng Lin, Cheng Yang 0016
ACM Multimedia6
2024 An End-to-End Real-World Camera Imaging Pipeline
abstract
pipeline still faces challenges including the lack of joint optimization in system components, computational redundancies, and optical distortions such as lens shading.In light of this, we propose an end-to-end camera imaging pipeline (RealCamNet) to enhance realworld camera imaging performance.Our methodology diverges from conventional, fragmented multi-stage image signal processing towards end-to-end architecture.This architecture facilitates joint optimization across the full pipeline and the restoration of coordinate-biased distortions.RealCamNet is designed for highquality conversion from RAW to RGB and compact image compression.Specifically, we deeply analyze coordinate-dependent optical distortions, e.g., vignetting and dark shading, and design a novel 2804
Kepeng Xu, Zijia Ma, Li Xu 0008, Gang He 0002, Yunsong Li 0001, Wenxin Yu 0001, Taichu Han, Cheng Yang 0016
ACM Multimedia6
2023 R2L-Net: Rapid Medical Image Segmentation Network Regularized by Self-Supervised Relative Localization Task
abstract
Medical image segmentation plays a crucial role in medical imaging, especially with advancements in techniques like magnetic resonance imaging (MRI) and computed tomography (CT). UNet, a widely used architecture, has shown promising results in medical image segmentation. Several variants based on UNet, and Transformer-based models, like TransUNet, have also exhibited potential for improving segmentation performance. However, these models often require substantial data and computational resources, making them less suitable for on-the-fly segmentation in medical scenarios. This paper proposes a fast medical image segmentation network called R2L-Net, which leverages self-supervised and supervised learning. R2L-Net introduces a self-supervised relative localization task as a regularization term during network training to enhance performance. Compared to UNeXt, our proposed R2L-Net achieves superior results on two public datasets (ISIC and BUSI), with an improved Intersection over Union (IoU) by 5.22 and 5.36, respectively. Moreover, R2L-Net offers several advantages over existing models, including a small number of parameters, low computational complexity, and fast image processing.
Wenxin Yu 0001, Jun Gong 0001
BIBM2
2023 An Efficient and Robust Algorithm for Common Path Pessimism Removal In Static Timing Analysis
abstract
Common path pessimism removal (CPPR) is an essential step in modern static timing analysis (STA) to avoid unnecessary circuit overdesign caused by extra pessimism. However, current CPPR approaches exhibit good performance yet poor scalability. This work proposes an efficient algorithm based on timing graph pruning to address the CPPR scalability issue. On average, on industry benchmarks from TAU 2014 CAD contest, our algorithm runs 23.1% faster versus OpenTimer, the state-of-the-art open-source STA engine in literature.
Mengshi Gong, Wenxin Yu 0001
ACM Great Lakes Symposium on VLSI3
2023 Image Inpainting with Semantic-Aware Transformer
abstract
Image inpainting has made huge strides benefiting from the advantages of convolutional neural networks (CNNs) in understanding high-level semantics. Recently, some studies have applied transformers to the visual field to solve the problem that the convolution kernel cannot attend to longdistance information. However, unlike other vision tasks, there is much interference from damaged information in image inpainting tasks. We propose a new Semantic-Aware Transformer, which in addition to including a self-attention block like previous vision transformers, also has a block for learning semantics from QSVM. Specifically, to provide more valid information, we design a Quantized Semantic Vector Memory (QSVM) that encodes and saves semantic features in images as quantized vectors in latent space. Experiments on different datasets demonstrate the effectiveness and superiority of our method compared with the existing state-of-the-art.
Wenxin Yu 0001, Qi Wang 0089, Jun Gong 0001
ICASSP2
2023 Boosting Transferability of Adversarial Example via an Enhanced Euler's Method
abstract
Adversarial examples are intentionally designed images to force convolution neural networks to give error classification outputs. Existing attacks have constructed transferable adversarial examples from the base attack algorithm, data augmentation, ensemble model, etc. Nevertheless, under the black-box case especially facing defense models, the transferability of adversarial examples still needs to be improved. In this paper, we try to develop a better base attack to boost the transferability of adversarial examples. Through analyzing the baseline gradient-based attacks, we found their iterative procedures of updating gradients are similar to numerical Euler’s methods. From the perspective of numerical analysis, we employ an enhanced Euler’s method, with less approximate errors and thus more accurate, to search a better approximate optimal solution to construct a more transferable gradient-based attack. To this end, we apply two-step gradient calculations of the enhanced Euler’s method to correct gradient descent directions. As a base attack, our attacks can be easily integrated with data augmentations and ensemble model augmentations. Experimental results show the proposed augmented attack significantly improves the transferability of adversarial examples and achieves an average attack success rate at least 3% higher than state-of-the-arts under black-box settings with defense mechanisms.
Anjie Peng, Hui Zeng 0002, Wenxin Yu 0001, Xiangui Kang
ICASSP4
2023 Clock Aware Low Power Placement
abstract
In modern VLSI design, more than 30% of power consumption is caused by clock networks due to their large capacitance demand and high switching frequency. Prior research on clock power optimization primarily focuses on improving the routing and synthesis processes, where the planning flexibility is restricted by the placed registers. In this paper, we develop a novel co-optimization framework to conduct analytic placement and clock tree synthesis simultaneously. A numerical engine is proposed to balance the power demand between clock and regular signal networks by effectively and efficiently updating the clock tree topology, generating the network synthesis solution, and aligning it with the placement objective using the ePlace infrastructure. The experimental results validate the high performance of our proposed algorithm on all eight CLKISPD05 benchmarks, achieving a 45.1% reduction in clock-net wirelength and a 12.7% reduction in total switching power compared to RePlAce. Moreover, our algorithm outperforms the state-of-the-art clock aware placement algorithm SimPL+Lopper, achieving a 25% reduction in clock-net wirelength and a 10.5% reduction in total switching power.
Jinghao Ding, Linhao Lu, Zhaoqi Fu, Mengshi Gong, Yuanrui Qi, Wenxin Yu 0001
ICCAD7
2023 An Enhanced Neuron Attribution-Based Attack Via Pixel Dropping
abstract
Convolutional neural networks (CNNs) are vulnerable to adversarial examples (AEs). Existing feature-level attacks explore the neuron importance to distort the intrinsic object-aware features which are shareable among different CNNs, thus achieving great performance in transferability. In this work, we propose an enhanced neuron attribution-based attack via pixel dropping (ENAA) and try to increase the number of positive neurons to distort the object-aware features more fully than NAA. Specifically, when computing neuron attribution, we use a pixel dropping scheme to expand the regions where the source model pays attention to the image. Our ENAA can make the target model shift the attention regions of AEs far away from those of clean images. Experimental results validate that the proposed method outperforms the state-of-the-art feature-level attacks both in white-box and black-box settings.
Anjie Peng, Hui Zeng 0002, Wenxin Yu 0001
ICIP5
2023 TSFC: Texture and Structure Features Coupling for Image Inpainting
abstract
Image inpainting has made significant progress benefiting from the advantages of convolutional neural networks (CNNs). Deep learning-based methods have shown extraordinary performance in this field. In this paper, we propose a novel image inpainting architecture with pure CNN that can jointly reconstruct the structure and texture of the image. Our generative network architecture (TSFC) consists of two parallel stages: structure generation and texture generation. In the structure generation stage, we use the large convolution kernel, which is highly neglected in modern networks, using the effective perceptual field of the large convolution kernel to enhance the perception of overall structural features. In the texture generation stage, we use the small convolution kernel to extract local texture features. Qualitative and quantitative experimental results on CelebA-HQ and Paris Street View datasets demonstrate the effectiveness and superiority of our method.
Qi Wang 0089, Wenxin Yu 0001, Jun Gong 0001
ICIP3
2023 BCKD: Block-Correlation Knowledge Distillation
abstract
In this paper, we propose Block-Correlation Knowledge Distillation (BCKD), a novel and efficient knowledge distillation method that differs from the classical method, using the simple multilayer-perceptron (MLP) and the classifier of the pre-trained teacher to train the correlations between adjacent blocks of the model. Over the past few years, the performance of some methods has been restricted by the feature map size or the lack of samples in small-scale datasets. By our proposed BCKD, the above problem is satisfactorily solved and has a superior performance without introducing additional overhead. Our method is validated on CIFAR100 and CI-FAR10 datasets, and experimental results demonstrate the effectiveness and superiority of our method.
Qi Wang 0089, Wenxin Yu 0001, Jun Gong 0001
ICIP3
2023 Dynamic Unilateral Dual Learning for Text to Image Synthesis
abstract
Dual learning trains two inverse processes tasks dually to further improve the selected tasks’ performance. There are currently two training paradigms in dual learning. One is to directly utilize two existing models for training in a dual manner to improve the selected models’ performance. However, it cannot effectively guarantee the improvement of the selected models. Another is that the networks of both parties are manually designed. Nevertheless, the network performance of both parties will be poor in the initial stage of training, which will easily lead to unsatisfactory training. Besides, most of dual learning researches can only be used for the conversion between the same data types, but it is powerless for the conversion of different data types. To address the above issues, a paradigm called unilateral dual learning (UDL) is proposed and verified in the text-to-image (T2I) synthesis field. UDL allows one party to design the network manually, and the other party calls the pre-trained model to promote the training of the manually designed network to achieve satisfactory training. Experimental results on the Oxford-102 flower and Caltech-UCSD Birds datasets demonstrate the feasibility of our proposed UDL paradigm in the T2I field, and it achieves excellent performance qualitatively and quantitatively.
Jiayao Xu, Ryugo Morita, Wenxin Yu 0001, Jinjia Zhou
ICIP4
2023 Text to Image Generation with Conformer-GAN
Zhiyu Deng, Wenxin Yu 0001, Lu Che, Jun Shang, Jun Gong 0001
ICONIP (5)2
2023 Fooling Downstream Classifiers via Attacking Contrastive Learning Pre-trained Models
Chenggang Li, Anjie Peng, Hui Zeng 0002, Wenxin Yu 0001
ICONIP (12)5
2023 Detecting Adversarial Examples via Classification Difference of a Robust Surrogate Model
Anjie Peng, Kang Deng, Hui Zeng 0002, Wenxin Yu 0001
ICONIP (12)5
2023 Text-to-Image Synthesis with Threshold-Equipped Matching-Aware GAN
Jun Shang, Wenxin Yu 0001, Lu Che, Hongjie Cai, Zhiyu Deng, Jun Gong 0001
ICONIP (12)2
2023 Neuron Attribution-Based Attacks Fooling Object Detectors
Guoqiang Shi, Anjie Peng, Hui Zeng 0002, Wenxin Yu 0001
ICONIP (13)4
2023 Multi-view Consistency View Synthesis
Xiaodi Wu 0005, Wenxin Yu 0001, Yufei Gao 0002, Jun Gong 0001
ICONIP (12)3
2023 Depth Normalized Stable View Synthesis
Xiaodi Wu 0005, Wenxin Yu 0001, Yufei Gao 0002, Jun Gong 0001
ICONIP (14)3
2023 Image Inpainting with Semantic U-Transformer
Lingfan Yuan, Wenxin Yu 0001, Lu Che
ICONIP (14)2
2023 Global priors guided modulation network for joint super-resolution and SDRTV-to-HDRTV
abstract
Watching low resolution standard dynamic range (LR SDR) video on a 4K high dynamic range (HDR) TV is not the best viewing experience. Joint super-resolution (SR) and SDRTV-to-HDRTV aims to enhance the visual quality of LR SDR videos that have quality deficiencies in resolution and dynamic range. Previous methods that rely on learning local information typically cannot do well in preserving color conformity and long-range structural similarity, resulting in unnatural color transition and texture artifacts. In order to tackle these challenges, we propose a global priors guided modulation network (GPGMNet). In particular, we design a global priors extraction module (GPEM) to extract color conformity prior and structural similarity prior that are beneficial for SDRTV-to-HDRTV and SR tasks, respectively. To further exploit the global priors and preserve spatial information, we devise multiple global priors-guided spatial-wise modulation blocks (GSMBs) with a few parameters for intermediate feature modulation. In these GSMBs, the modulation parameters are generated by the shared global priors and the spatial features map from the spatial pyramid convolution block (SPCB). With these elaborate designs, the GPGMNet can achieve higher visual quality with lower computational complexity. Extensive experiments demonstrate that our proposed GPGMNet is superior to the state-of-the-art methods. Specifically, our proposed model exceeds the state-of-the-art by 0.64 dB in PSNR, with 69% fewer parameters and 3.1 × speedup.
Gang He 0002, Shaoyi Long, Li Xu 0008, Chang Wu 0001, Wenxin Yu 0001, Jinjia Zhou
Neurocomputing5
2023 Positive-Unlabeled Learning for Knowledge Distillation
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001
Neural Process. Lett.3
2022 An Efficient Maze Routing Algorithm for Fast Global Routing
abstract
Maze routing remains the most time-consuming step for modern global routers. Previous works accelerate the maze routing by routing multiple regions or nets simultaneously. This paper presents a novel parallel maze router with bidirectional path search and dynamic routing scheduling, which exhibits higher efficiency than all the previous routers. On the ISPD 2008 benchmark suite, our router outperforms the fastest global routers SPRoute and FastRoute 4.1 by an average speedup of 1.95x and 10.03x, while the difference on the total overflow and wirelength is negligible.
Zhaoqi Fu, Wenxin Yu 0001, Xin Cheng 0004
ACM Great Lakes Symposium on VLSI2
2022 Optimal Region-based Mixed-Cell-Height Detailed Placement Considering Complex Minimum-Implant-Area Constraints
abstract
We propose a minimum-implant-area (MIA) aware detailed placement algorithm for multi-row-height standard cells. Specifically, (1) we calculate the optimal regions for all the cells, and (2) we cluster each group of cells with the same threshold voltage and width less than the minimum implant width and then reshape each cluster. (4) We develop an enhanced legalization algorithm to minimize the total wirelength. (5) We solve the remaining inter-row violations by greedily shifting the concerning cells with minimum displacement. Compared with the state-of-the-art work [6], the experimental results show that on average of all the ISPD 2014 benchmarks A [1] our algorithm reduces the wirelength by 6% and runs 6.02x faster with all the MIA violations resolved.
Wenxin Yu 0001, Zhaoqi Fu, Xin Cheng 0004
ACM Great Lakes Symposium on VLSI2
2022 An Enhanced Transferable Adversarial Attack of Scale-Invariant Methods
abstract
Scale-invariant method (SIM) is a state-of-the-art model augmentation method to improve the transferability of adversarial examples. However, we find that SIM is easily affected by the scaling operation with small scaling factors, and cannot stably enhance the transferability of the base attack. In this paper, we propose an enhanced transferable attack based on SIM. To alleviate the instability of SIM caused by the scaled copy which does not satisfy scale-invariance, we propose to ensemble logit-outputs of scale copies of the input image, rather than ensemble the gradients, to form an ensemble attack that generates transferable adversarial images from multiple models of the original CNN model. Compared with the existing ensemble methods, our method is fast yet effective and can be easily integrated into the gradient-based attacks. The experimental results show that the proposed integrated EL-NI-FGSM attack stably improves the transferability of NI-FGSM, and outperforms SI-NI-FGSM, achieving >8% higher of attack success rate for both white-box and black-box attacks on CIFAR-10.
Anjie Peng, Rong Wei, Wenxin Yu 0001, Hui Zeng 0002
ICIP4
2022 Text-Guided Image Manipulation Based on Sentence-Aware and Word-Aware Network
abstract
Text-guided image manipulation aims to use the given text description to modify the semantic content of the corresponding part in the input image. Although researchers have been obtained satisfactory performance in this field, they only 1) utilize the global sentence information at the initial modification stage and 2) exploit the fixed word information for regional adjustment in the subsequent modification process, hindering the improvement of image manipulation quality. Motivated by the mentioned issues, this paper proposes a novel approach to improve the performance of text-guided image manipulation by using sentence-aware and word-aware network. Concretely, we utilize global sentence information throughout the image manipulation process to improve the semantic consistency with the input text. On the other hand, we employ the dynamic selection method to dynamically adjust the word information corresponding to the regional image content to further improve the manipulation quality. As a result, our work surpasses the existing state-of-the-art methods on CUB and Oxford-102 flower datasets, demonstrating our effectiveness and superiority. In terms of Inception Score, our proposed method performs the most excellent performance. In terms of NIMA, the score of our method is closest to the score of the original dataset images, proving that our manipulated results are the most authentic.
Man M. Ho, Jinjia Zhou, Ning Jiang 0002, Wenxin Yu 0001
ICME6
2022 Countering the Anti-detection Adversarial Attacks
Anjie Peng, Chenggang Li, Xiaofang Huang, Hui Zeng 0002, Wenxin Yu 0001
ICONIP (4)6
2022 Effect of Image Down-sampling on Detection of Adversarial Examples
Anjie Peng, Chenggang Li, Hui Zeng 0002, Wenxin Yu 0001
ICONIP (4)7
2022 FPD: Feature Pyramid Knowledge Distillation
Qi Wang 0089, Wenxin Yu 0001, Xuewen Zhang, Jun Gong 0001
ICONIP (1)3
2022 A Point Matching Strategy of 3D Loss Function for Single RGB Images Deep Mesh Reconstruction
abstract
Recent the-state-of-the-art image-based three-dimensional (3D) reconstruction methods that represent 3D shapes mainly using triangular mesh because of its memory efficiency and ability to present surface detail of objects compared to voxel and point cloud. Previous works usually follow an encoding and decoding pattern. A deep neural network to extract the features from the picture and reconstruct the 3D structure. It is a typical supervised learning process, requiring loss function to supervise the training. No existing works directly calculate the loss between the reconstruction mesh and ground truth mesh. Instead, they indirectly used the Chamfer Distance (CD) between point clouds as the loss. Most of the previous works focus on the encoding and decoding parts instead of the loss and CD is used for all works. However, when CD is applied to two point clouds with the same number of points, some points can match any number of points in another point cloud, so some points will be less involved in calculating the loss function, which will reduce the utilization of information. Therefore, We propose a new point matching strategy to calculate the loss. The point matching strategy we proposed limits the maximum number of matches for each point, allowing more points to be more involved in the loss calculation, thereby improving the information utilization rate. Experiments on single view reconstruction (SVR) and auto-encoding methods show that this new loss method can replace CD in this type of works and has better training results and 3D reconstruction quality.
Ning Jiang 0002, Jiarui Cheng, Yufei Gao 0002, Wenxin Yu 0001
ISCAS6
2022 Music to Dance: Motion Generation Based on Multi-Feature Fusion Strategy
abstract
Studies on generating dance sequences from music can greatly promote dance teaching and animation production. However, the current results of such tasks are not very satisfactory. Due to the lack of consideration of human body structure, unreal phenomena such as slippery feet and twisted joints will appear in the resulting movements. Since only a single correlation feature between music and dance is learned, there is a problem of poor matching between generated actions and music. To solve these problems, we propose a learning framework based on GAN and a multi-feature fusion strategy to realize the mapping from music to dance. We use two discriminators to constrain style and authenticity respectively to make our actions generated by the network more coherent and natural, which is essential in producing authentic dance sequences. Moreover, we extracted the three characteristics of music style, beat, and structure. Then we used the feature fusion method to obtain a comprehensive feature representation, which guarantees the consistency of the generated dance sequence with the input music to the greatest extent. The experimental quantitative and qualitative results show that our method can generate accurate, consistent, beat-matching dance movements from music. In addition, we also collected and produced a data set containing music features and corresponding pose sequences, which is convenient for pose generation based on music features.
Yufei Gao 0002, Wenxin Yu 0001, Xuewen Zhang
ISCAS2
2022 Improve 3D Feature Extraction and Fusion for Stage Diagnosis of Alzheimer's Disease
abstract
Alzheimer’s disease (AD) is typical dementia, which is progressive and irreversible. Usually, the clinical diagnosis of patients is at a later stage, so early diagnosis can control the patient’s condition in time. The doctor is usually diagnosing the patient’s condition by 3D brain magnetic resonance imaging (MRI). However, because the 3D MRI structures of adjacent stages are almost similar, the multi-class diagnosis of AD becomes difficult. Therefore, there is a need to enhance the ability to extract more discriminative features from 3DMRI, promoting more accurate diagnosis. In addition, not only the entire MRI changes but also local areas in the MRI. Therefore, it is necessary to pay attention to changes in the entire image and local areas and to fuse features of different scales. In this paper, we propose an innovative convolutional network architecture for feature extraction and feature fusion. It consists of three modules: 1) a network based on ResNet-10, 2) 3D asymmetric convolution block (ACB), 3) multi-scale channel attentional feature fusion (MS-CAFF) module. The proposed model has been tested on the ADNI dataset and achieved an accuracy of 88.33%, which is nearly 2% higher than the latest research.
Mingjin Liu, Wenxin Yu 0001, Jialiang Tang, Ning Jiang 0002, Kang Xu 0002
ISCAS2
2022 Text-to-image synthesis: Starting composite from the foreground content
Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002
Inf. Sci.3
2021 A Progressive Image Inpainting Algorithm with a Mask Auto-update Branch
Liang Nie, Wenxin Yu 0001, Xuewen Zhang, Siyuan Li 0004, Ning Jiang 0002
ICANN (2)2
2021 Drawgan: Text to Image Synthesis with Drawing Generative Adversarial Networks
abstract
In this paper, we propose a novel drawing generative adversarial networks (DrawGAN) for text-to-image synthesis. The whole model divides the image synthesis into three stages by imitating the process of drawing. The first stage synthesizes the simple contour image based on the text description, the second stage generates the foreground image with detailed information, and the third stage synthesizes the final result. Through the step by step synthesis process from simple to complex and easy to difficult, the model can draw the corresponding results step by step and finally achieve the higher-quality image synthesis effect. Our method is validated on the Caltech-UCSD Birds 200 (CUB) dataset and the Microsoft Common Objects in Context (MS COCO) dataset. The experimental results demonstrate the effectiveness and superiority of our method. In terms of both subjective and objective evaluation, our method’s results surpass the existing state-of-the-art methods.
Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002
ICASSP3
2021 Gradient Local Binary Pattern For Convolutional Neural Networks
abstract
Convolutional neural networks(CNNs) have achieved a performance significantly superior to traditional machine learning methods. However, in the traditional machine learning methods, the feature extraction algorithms are compelling and beneficial for CNNs. This paper introduces the classic feature extraction algorithm gradient local binary pattern(GLBP) to the CNNs. More specially, the GLBP extractor weights will be fixed into the $3\times 3$ sized kernels to construct the GLBP layer to replace the first layer of CNNs. In the GLBP layer, the features extracted by the GLBP kernels will concate or add to the feature process by the convolutional kernels. Through extensive experiments, we demonstrated that the GLBP layer could efficiently improve CNNs performance. When training on the ImageNet dataset, the ResNet18 with GLBP layer obtained 1.19% Top-1 accuracy improvement and 0.87% Top-5 accuracy improvement, respectively.
Jialiang Tang, Ning Jiang 0002, Wenxin Yu 0001
ICIP3
2021 Text To Image Synthesis With Erudite Generative Adversarial Networks
abstract
In this paper, an Erudite Generative Adversarial Networks (EruditeGAN) is proposed for the text-to-image synthesis task. By introducing additional image distribution related to the original image into the network structure, the entire network can learn more about the image distribution and become more knowledgeable. In this case, it can be more clear about the distribution of the image that needs to be synthesized and finally synthesize high-quality results. Experiments well validate our method’s effectiveness and demonstrate the different effects of different distribution situations on the final results. According to the quantitative results of Fréchet Inception Distance (FID) and R-precision, our method’s comprehensive score is the best, which reflects our results are closer to the real image effect.
Wenxin Yu 0001, Ning Jiang 0002, Jinjia Zhou
ICIP2
2021 SCAN: Spatial and Channel Attention Normalization for Image Inpainting
Wenxin Yu 0001, Liang Nie, Xuewen Zhang, Siyuan Li 0004, Jun Gong 0001
ICONIP (6)2
2021 Using a Two-Stage GAN to Learn Image Degradation for Image Super-Resolution
Jiarui Cheng, Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001
ICONIP (5)5
2021 Improving Shallow Neural Networks via Local and Global Normalization
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001
ICONIP (1)4
2021 Attention-Based 3D ResNet for Detection of Alzheimer's Disease Process
Mingjin Liu, Jialiang Tang, Wenxin Yu 0001, Ning Jiang 0002
ICONIP (1)3
2021 Progressive Inpainting Strategy with Partial Convolutions Generative Networks (PPCGN)
Liang Nie, Wenxin Yu 0001, Siyuan Li 0004, Ning Jiang 0002, Xuewen Zhang, Jun Gong 0001
ICONIP (6)2
2021 Free-Form Image Inpainting with Separable Gate Encoder-Decoder Network
Liang Nie, Wenxin Yu 0001, Xuewen Zhang, Siyuan Li 0004, Jun Gong 0001
ICONIP (3)2
2021 Data-Free Knowledge Distillation with Positive-Unlabeled Learning
Jialiang Tang, Xin Cheng 0004, Ning Jiang 0002, Wenxin Yu 0001
ICONIP (2)5
2021 Consistent Knowledge Distillation Based on Siamese Networks
Jialiang Tang, Xin Cheng 0004, Ning Jiang 0002, Wenxin Yu 0001
ICONIP (5)5
2021 Transformer with Prior Language Knowledge for Image Captioning
Daisong Yan, Wenxin Yu 0001, Jun Gong 0001
ICONIP (2)2
2021 Triplet Mapping for Continuously Knowledge Distillation
Jialiang Tang, Ning Jiang 0002, Wenxin Yu 0001
ICONIP (1)4
2021 Deep Learning Based Placement Acceleration for 3D-ICs
Wenxin Yu 0001, Xin Cheng 0004, Jun Gong 0001
ICONIP (5)1
2021 QS-Hyper: A Quality-Sensitive Hyper Network for the No-Reference Image Quality Assessment
Xuewen Zhang, Yunye Zhang, Wenxin Yu 0001, Liang Nie, Ning Jiang 0002, Jun Gong 0001
ICONIP (4)3
2021 Semi-supervised Learning with Conditional GANs for Blind Generated Image Quality Assessment
Xuewen Zhang, Yunye Zhang, Wenxin Yu 0001, Liang Nie, Jun Gong 0001
ICONIP (3)3
2021 Triplet Knowledge Distillation Networks for Model Compression
abstract
Knowledge distillation is a widely used neural network model compression technique. In general, the knowledge distillation transfer the knowledge from a large pre-trained teacher network with superior performance to a small student network enables the student network to achieve better performance. This paper proposes a triplet knowledge distillation framework (abbreviated as TKD), which introduces a smaller assistant network into the knowledge distillation structure. The performance of the assistant network is lower than that of the student network. During the training of the TKD, by minimizing the Mean Squared Error(MSE) loss function, the output of the student network will closer to the output of the teacher network and further from that of the assistant network. Therefore, the student network can learn more expressive knowledge from the teacher network while throwing away mistaken knowledge in the assistant network. Finally, the student network achieves a surprising performance even superior to the teacher network. We have demonstrated the effectiveness of TKD by extensive experiments on benchmark datasets(CIFAR-10, CIFAR-100, SVHN, STL-10). When using VGGNet as an experimental model, the student network VGGNet13 achieving 94.29%, 75.30%, 95.53%, and 87.61% accuracy on the CIFAR-10, CIFAR-100, SVHN, and STL-10 datasets, improved by 1.24%, 2.81%, 0.40%, and 2.32%, respectively.
Jialiang Tang, Ning Jiang 0002, Wenxin Yu 0001, Wenqin Wu
IJCNN3
2021 Text to Image Synthesis based on Multi - Perspective Fusion
abstract
In this paper, we propose a multi-perspective fusion method to improve the performance of text-to-image synthesis. From the perspective of the generator, we introduce a dynamic selection method to make the text feature match the corresponding image feature better, while the multi-class discriminant method with mask segmentation image as the extra type is introduced from the perspective of the discriminator to improve its discrimination ability. Through the effective integration of these two aspects of improvement, more excellent results by our method are obtained. Experiments on the Caltech-UCSD Birds 200 (CUB) and Microsoft Common Objects in Context (MS COCO) datasets demonstrate our method's effectiveness and superiority. The qualitative and quantitative experiments validate that our method is superior to the existing state-of-the-art methods.
Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002
IJCNN4
2021 Gradient Local Binary Pattern Layer to Initialize the Convolutional Neural Networks
abstract
Deep neural network technology is a milestone achievement in the field of computer vision. It obtained the performance that the shallow network cannot achieve through the multi-layer network structure and the learning method of reverse adjustment parameters. However, the feature extraction algorithm of the shallow network is very effective and also is more beneficial for deep neural networks. In this paper, we combine the shallow network algorithm to proposes the gradient local binary pattern layer(GLBP layer) to replace the first layer of Convolutional Neural Networks(CNNs). The GLBP layer plays a role in initializing the CNNs and can improve network performance without increasing the number and complexity of network layers. In the experiment, using the extracted layer modified by the GLBP feature algorithm to replace other classic deep neural networks, 2.65% and 2.9% performance improvements were obtained in the WideResNet16-2 and ResNet-101 respectively when training on CIFAR-100 dataset.
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou, Liuwei Mai
ISCAS3
2021 Data-Free Network Pruning for Model Compression
abstract
Convolutional neural networks(CNNs) are often over-parameterized and cannot apply to existing resource-limited artificial intelligence(AI) devices. Some methods are proposed to model compress the CNNs, but these methods are data-driven and often unable when lacking data. To solve this problem, in this paper, we propose a data-free model compression and acceleration method based on generative adversarial networks and network pruning(named DFNP), which can train a compact neural network only needs a pre-trained neural network. The DFNP consists of the source network, generator, and target network. First, the generator will generate the pseudo data under the supervise of the source network. Then the target network will get by pruning the source network and use these generated data for training. And the source network will transfer knowledge to the target network to promote the target network to achieve a similar performance of the source network. When the VGGNet- 19 is select as the source network, the target network trained by DFNP contains only 25% parameters and 65% calculations of the source network. Still, it retains 99.4% accuracy on the CIFAR-10 dataset without any real data.
Jialiang Tang, Mingjin Liu, Ning Jiang 0002, Huan Cai, Wenxin Yu 0001, Jinjia Zhou
ISCAS5
2021 Spatial and Channel Dimensions Attention Feature Transfer for Better Convolutional Neural Networks
abstract
Knowledge distillation is an extensively researched model compression technology, which uses a large teacher network to transmit information to a small student network. The critical point of the knowledge distillation method to improve the performance of the student network is to find an effective method to extract the information from the feature. The attention mechanism is a widely used feature processing method to process features effectively and obtain more expressive information. In this paper, we propose to use the dual attention mechanism in knowledge distillation to improve the performance of student networks, which extracts information from the spatial and channel dimensions of the feature. The channel dimension attention is search 'what' channel is more meaningful, and the spatial dimension attention is determine 'where' part of the feature is more expressive in a feature map. We have conducted extensive experiments on different datasets, shown that by implementing a dual attention mechanism to extract more expressive information for knowledge transfer, the student network can achieve performance beyond the teacher network.
Jialiang Tang, Mingjin Liu, Ning Jiang 0002, Wenxin Yu 0001, Changzheng Yang
ISCAS4
2021 Knowledge Distillation Based on Positive-Unlabeled Classification and Attention Mechanism
abstract
With the rapid development of deep learning, convolutional neural networks(CNNs) have achieved great success. But these high-capability CNNs often with a huge burden of computation and memory, which hinders these CNNs from applying to practical application. To solve this problem, in this paper, we proposed a method to train a compact model with high-capacity. The student network with fewer parameters and calculations will learning from the knowledge of the teacher network with more parameters and calculations. To promote the ability of the student network, the more expressive knowledge is extracted from the middle-layer feature of neural networks by attention mechanism, and the knowledge transforms more effective from the teacher network to the student network by the positive- unlabeled(PU) classifier. We validate our method in extensive experiments, showing that it can train the student network to achieve significant performance superior to the teacher network.
Jialiang Tang, Mingjin Liu, Ning Jiang 0002, Wenxin Yu 0001, Changzheng Yang, Jinjia Zhou
ISCAS4
2021 SPS: A Subjective Perception Score for Text-to-Image Synthesis
abstract
A fundamental problem of text-to-image synthesis is the lack of quality assessment for a single generated image. Quantitative indicators of this work (such as Inception Score and Fréchet Inception Distance) only affect plenty of images' feature distribution. It causes monotonous evaluation and plenty of poor-quality image results. This paper proposes a new evaluation criterion for text-to-image synthesis by the blind image quality assessment(BIQA) method. To train the model, a Multi-Metrics Quality Assessment Dataset for generated birds' images(MMQA) is proposed. Besides, the Multi-hyper model is proposed to fit our dataset better. Experiments show that our method evaluates text- to-image tasks more comprehensively and optimize their results.
Xuewen Zhang, Wenxin Yu 0001, Ning Jiang 0002, Yunye Zhang
ISCAS2
2021 Local Feature Normalization
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou
KSEM3
2020 Interactive Separation Network For Image Inpainting
abstract
Image inpainting, also known as image completion, is the process of filling in the missing region of an incomplete image to make the repaired image visually plausible. Strided convolutional layer learns high-level representations while reducing the computational complexity, but fails to preserve existing detail from the original images (eg, texture, sharp transients), therefore it degrades the generative model in image inpainting task. To reduce the erosion of high-resolution components of images meanwhile maintaining the semantic representation, this paper designs a brand-new network called Interactive Separation Network that progressively decomposites the features into two streams and fuses them. Besides, the rationality of network design and the efficiency of proposed network is demonstrated in the ablation study. To the best of our knowledge, the experimental results of proposed method are superior to state-of-the-art inpainting approaches.
Siyuan Li 0004, Xin Cheng 0004, Kepeng Xu, Wenxin Yu 0001, Gang He 0002, Jinjia Zhou
ICIP6
2020 Gpu Accelerated Polar Fourier Analysis For Feature Extraction
abstract
Polar Fourier analysis can extract orthogonal and rotation invariant features and demonstrate superior performances in many image processing tasks. For real world applications, execution efficiency is always a significant challenge. With widespread use of Graphics Processing Unit (GPU), this research presents GPU accelerated polar Fourier analysis. Proposed parallel algorithm is based on mathematical properties of polar Fourier analysis and optimization techniques of GPU. Optimal parameter selections for GPU execution are also evaluated. In our experiments, with same computation result proposed method is over 170 times faster. Wide range of applications that using polar Fourier analysis will inspired from this work.
Mingkai Tang 0002, Zhuozhang Li, Yinwei Zhan, Wenxin Yu 0001
ICIP6
2020 Coarse-to-Fine Attention Network via Opinion Approximate Representation for Aspect-Level Sentiment Classification
Wei Chen 0062, Wenxin Yu 0001, Gang He 0002, Ning Jiang 0002, Gang He 0001
ICONIP (1)2
2020 Search-and-Train: Two-Stage Model Compression and Acceleration
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou
ICONIP (5)4
2020 LPI-Net: Lightweight Inpainting Network with Pyramidal Hierarchy
Siyuan Li 0004, Kepeng Xu, Wenxin Yu 0001, Ning Jiang 0002
ICONIP (4)4
2020 Gradient-Based Adversarial Image Forensics
Anjie Peng, Kang Deng, Shenghai Luo, Hui Zeng 0002, Wenxin Yu 0001
ICONIP (2)6
2020 Customizable GAN: Customizable Image Synthesis Based on Adversarial Learning
Wenxin Yu 0001, Jinjia Zhou, Xuewen Zhang, Jialiang Tang, Siyuan Li 0004, Ning Jiang 0002, Gang He 0001, Gang He 0002
ICONIP (4)2
2020 No-Reference Quality Assessment Based on Spatial Statistic for Generated Images
Yunye Zhang, Xuewen Zhang, Wenxin Yu 0001, Ning Jiang 0002, Gang He 0002
ICONIP (4)4
2020 Deep Feature Compatibility for Generated Images Quality Assessment
Xuewen Zhang, Yunye Zhang, Wenxin Yu 0001, Ning Jiang 0002, Gang He 0001
ICONIP (4)4
2019 Text to Image Synthesis Based on Multiple Discrimination
Yunye Zhang, Wenxin Yu 0001, Jingwei Lu, Li Nie, Gang He 0001, Ning Jiang 0002, Gang He 0002, Yibo Fan
ICANN (3)3
2019 Accelerated Detail-Enhanced Ambient Occlusion
abstract
Ambient Occlusion (AO) is a technique to approximate the effect of environment lighting and add realism to a scene by accentuating surface details and adding soft shadows, which is widely used in multimedia applications. Neural Network Ambient Occlusion (NNAO) is a pioneer in introducing deep learning to accurate and real-time AO, but it has two limitations: 1) performance bottleneck under excessive amount of samples ; 2) low contrast and blurred edges leading to unreal effect. To overcome these two limitations, we propose Accelerated Detail-enhanced Ambient Occlusion (ADAO) method based on NNAO by adopting three image processing methods: 1) spiral sampling in screen space; 2) contrast enhancement of AO map; 3) normal-depth edge preserving bilateral filtering. Experimental results show that the proposed method is over 2 times faster than NNAO and produces shadows with more realistic details.
Yinwei Zhan, Jujian Lv, Qieshi Zhang, Wenxin Yu 0001
ICIP6
2019 Target-Based Attention Model for Aspect-Level Sentiment Analysis
Wei Chen 0062, Wenxin Yu 0001, Yunye Zhang, Kepeng Xu, Fengwei Zhang, Yibo Fan, Gang He 0002
ICONIP (3)2
2019 Inpainting with Sketch Reconstruction and Comprehensive Feature Selection
Siyuan Li 0004, Zhijing Li 0003, Kepeng Xu, Matthieu Claisse, Wenxin Yu 0001, Gang He 0001, Gang He 0002, Yibo Fan
ICONIP (5)6
2019 IRSNET: An Inception-Resnet Feature Reconstruction Model for Building Segmentation
Kepeng Xu, Li Nie, Wenxin Yu 0001, Yunye Zhang, Wei Chen 0062, Siyuan Li 0004, Shangwei Deng, Yibo Fan, Hui Zhang 0051, Valentin Bouillon
ICONIP (5)4
2019 Dense Image Captioning Based on Precise Feature Extraction
Yunye Zhang, Wenxin Yu 0001, Li Nie, Gang He 0002, Yibo Fan
ICONIP (5)4
2019 Text to Image Synthesis Using Two-Stage Generation and Two-Stage Discrimination
Yunye Zhang, Wenxin Yu 0001, Gang He 0001, Ning Jiang 0002, Gang He 0002, Yibo Fan
KSEM (2)3
2019 Maximum Satisfiability Formulation for Optimal Scheduling in Overloaded Real-Time Systems
Xiaojuan Liao, Hui Zhang 0051, Miyuki Koshimura, Rong Huang 0003, Wenxin Yu 0001
PRICAI (1)5
2018 Q Value-Based Dynamic Programming with Boltzmann Distribution by Using Neural Network
Wenxin Yu 0001, Gang He 0001, Yibo Fan, Gang He 0002, Jiu Xu
ICONIP (7)1
2018 Perceptual model optimized efficient foveated rendering
abstract
Higher resolution, wider FOV and increasing frame rate of HMD are demanding more VR computing resources. Foveated rendering is a key solution to these challenges. This paper introduces a perceptual model optimized foveated rendering. Tessellation levels and culling areas are adaptively adjusted based on visual sensitivity. We improve rendering performance while satisfying visual perception.
Zipeng Zheng, Yinwei Zhan, Wenxin Yu 0001
VRST5
2017 Fast mode decision and PU size decision algorithm for intra depth coding in 3D-HEVC
Gang He 0002, Jing Hu 0005, Yunsong Li 0001, Wenxin Yu 0001, Peikun Liu, Ruixue Guo
J. Vis. Commun. Image Represent.4
2016 Fast algorithm based on sole- and multi-depth measurements for HEVC intra coding
abstract
In High Efficiency Video Coding (HEVC), intra coding plays an important role, but also involves huge computational complexity due to a flexible coding unit (CU) structure and a large number of prediction modes. This paper presents a fast algorithm based on the sole- and multi-depth measurements to reduce the complexity from CU and prediction mode decisions. For the CU decision, evaluation results with sole and multiple depths are utilized to judge if the CU is a heterogeneous, homogeneous, or depth prominent one, where fast CU decisions are made. For the prediction mode decision, the tendencies for different CU sizes are detected based on multiple depths. The number of searching modes is decreased adaptively for the depth with fewer tendencies. Experimental results show the proposed algorithm reduces 61.49% computational complexity, with 0.75% bit-rate increasing, which is more efficient than state-of-the-arts.
Gang He 0002, Jing Hu 0005, Yunsong Li 0001, Wenxin Yu 0001
ICIP4
2015 Frame compatible format fast encoder with stereo matching
abstract
In this paper, a frame compatible format fast encoder with stereo matching is proposed. Through packing the two neighboring views after down-sampling into one frame, the frame compatible format coding allows stereo video to be encoded on the conventional video applications. The matching of content similarity is the key of the fast encoder of Frame Compatible Format (FCF) H.264. Experiments show that the stereo matching can achieve a better result than simply using a shift obtaining method. This paper gives a statistical analysis of the prediction correlation between the two packed views in FCF by using different method to prove that the system can obtain much better results with stereo matching method. A modified fast stereo matching algorithm which is suitable for FCF fast encoding is introduced. Finally, the process and simulated results of stereo matching combined frame compatible format fast encoder is shown. Through the experimental result, the proposed algorithm can obtain a better result than the previous work. It can achieve better video quality, less PSNR lost and the lower bit rate increment.
Wenxin Yu 0001, Weichen Wang 0005, Jiu Xu
ISCAS1
2013 Combined hole-filling with spatial and temporal prediction
abstract
A combined hole-filling approach with spatial and temporal prediction is presented in this paper. Depth image-based rendering (DIBR) is generally used to synthesize virtual view images in free viewpoint television (FTV) and three-dimensional (3-D) video. Limited original camera views and depth maps are used to generate the additional virtual views in the synthesizing process. One of the main problems in DIBR is that there are some regions are occluded by the foreground objects in the original views, and they will be some holes in the generated additional virtual views, especially for the view extrapolation cases. Therefore, the proposed algorithm is introduced and it can be used to fill the holes which caused by disocclusion regions and inaccurate depth values. The proposed algorithm combines the spatial and temporal prediction, and the performance is much better and more stable than the previous work. The experimental results show that the proposed method can improve the quality of the virtual views a lot compared with the previous work. The improvement is not only obvious in the objective comparison, but also in the subjective comparison.
Wenxin Yu 0001, Weichen Wang 0005, Gang He 0002, Satoshi Goto
ICIP1
2013 Gradient Local Binary Patterns for human detection
abstract
In recent years, local pattern based features have attracted increasing interest in object detection and recognition systems. Local Binary Pattern (LBP) feature is widely used in texture classification and face detection. But the original definition of LBP is not suitable for human detection. In this paper, we propose a novel feature set named gradient local binary patterns (GLBP), Original GLBP and Improved GLBP, for human detection. Experiments are performed on INRIA dataset, which shows the proposal GLBP feature is more discriminative than histogram of orientated gradient (HOG), histogram of template (HOT) and Semantic Local Binary Patterns (S-LBP), under the same training method. In our experiments, the window size is fixed. That means the performance can be improved by boosting and cascade methods. And the computation of GLBP feature is parallel, which make it easy for hardware acceleration. These factors make GLBP feature possible for real-time human detection.
Ning Jiang 0002, Jiu Xu, Wenxin Yu 0001, Satoshi Goto
ISCAS3
2013 Multi-scale bidirectional local template patterns for real-time human detection
abstract
In this paper, a feature named multi-scale bidirectional local template patterns (MBLTP) is proposed for human detection. As an extension of bidirectional local template patterns (BLTP), MBLTP not only integrates the textural and gradient information according to the four predefined templates but also calculates information for additional feature vectors by adjusting the scale of the training samples. These additional feature vectors contain multi-scale information on the samples, which can make the feature more discriminative than its original form. Experimental results for an INRIA dataset show that the detection rate of our proposed MBLTP feature outperforms those of other features such as the multi-level histogram of orientated gradient (multi-level HOG), multi scale block histogram of template (MB-HOT), and HOG-LBP. Moreover, in order to make our feature meet real-time requirements, an implementation based on a graphic process unit (GPU) is adopted to accelerate the calculation.
Jiu Xu, Ning Jiang 0002, Xinwei Xue, Heming Sun, Wenxin Yu 0001, Satoshi Goto
MMSP5