EDBT 2026 Demo / reviewers in the wild / expert
Ding Liu 0001
dblp:26/5200-1
· DBLP profile ↗
51ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0002-0931-1345ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fairness-Aware influence blocking maximization: A centrality-Enhanced adversarial graph embedding framework
Lan Yang 0007, Ding Liu 0001, Zhiwu Li 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Verification and Enforcement of Concealability and Diagnosability in Probabilistic Timed AutomataabstractInternational audience Tareq Ahmad Al-Sarayrah, Sjood Ammen Daje, Ding Liu 0001, Mohamed Ghazel, Zhiwu Li 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Temporal Liveness Enforcement via Parameter Tuning in Dual-Time Petri Nets
Ruotian Liu, Yufeng Chen 0001, Maria Pia Fanti, Ding Liu 0001, Boyu Dong |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Identification of Deletion Attacks in Discrete-Event Systems Through Petri Nets Enhanced With an Observation Structure
Adeeb A. Ahmed, Yufeng Chen 0001, Ding Liu 0001, Zhiwu Li 0001 |
IEEE Trans. Reliab. | 3 |
| 2025 | Diagnosability Verification and Enforcement for Unbounded Petri Nets by Online SupervisorsabstractThis paper addresses the problems of diagnosability verification and enforcement of discrete event systems modeled with unbounded Petri nets. Diagnosability in such systems is critical for ensuring reliability and maintaining operational integrity, yet current methods often struggle with the complexity introduced by unboundedness and potential deadlocks. Given an unbounded labeled Petri net that may reach deadlocks, a quiescent basis coverability graph is established to verify the diagnosability of the considered system. This procedure employs a deterministic finite state automaton, called an extended verifier, derived from the proposed quiescent basis coverability graph. It is shown that an unbounded Petri net is diagnosable if and only if the verifier does not contain a class of cycles, called repetitive F-cycles. This result also provides necessary and sufficient conditions for diagnosability enforcement by developing an online supervisor. Further, the designed supervisor is maximally permissive and also circumvents a plant entering deadlocks by firing non-fault sequences. Examples are presented to demonstrate the proposed method. Note to Practitioners—Fault diagnosis and diagnosability enforcement are critical for the development and operation of highly automated systems covering computer-integrated production processes, intelligent traffic, computer and communication networks, smart gird, etc. This work touches upon this problem from the perspective of discrete event systems that are modeled with unbounded labeled Petri nets. The feasibility and applicability of the reported method stem from the usage of a structurally compact representation of a considered plant such that the computational cost of a real-world system is acceptable. The graphical representation of Petri nets as well as the proposed quiescent basis coverability graph make the method easy to use and manipulate. Moreover the sufficient and necessary conditions of diagnosability enforcement can be readily verified by the supervisory theory, facilitating its adoption by practitioners. Shaopeng Hu 0001, Yihui Hu, Ding Liu 0001, Maria Pia Fanti, Zhiwu Li 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Fairness-Aware Influence Blocking Maximization in Social NetworksabstractNowadays, social media continually influences people’s opinions and behaviors. The spread of unchecked rumors has a detrimental impact on society. It is crucial to develop effective strategies to block the dissemination of rumors as much as possible. However, current approaches for influence blocking overlook the fairness among diverse communities, which are fundamental structures in social networks, leading to a disproportionate exclusion of marginalized communities from intervention benefits and resulting in fairness disparities. This article addresses such a critical issue in influence blocking and proposes a fairness-aware blocking strategy that balances the effectiveness of influence blocking and fairness. A sampling-based approximate algorithm called fairness-aware influence blocking maximization based on martingales (FIBMM), which is applicable to large networks, is presented. We conduct a series of experiments on six real networks, showing that the FIBMM can effectively reduce fairness disparities between diverse communities while maintaining blocking effectiveness. Ding Liu 0001, Lan Yang 0007, Zhiwu Li 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | SU3plus: An Enhanced Swin-UNet for Synthesizing Vertically Integrated Liquid From Multiple Meteorological Satellite DataabstractRadar composite reflectivity is a crucial component of Earth observation data, playing a significant role in applications such as weather forecasting and climate disaster tracking. Due to the deployment challenges and limited coverage of meteorological radars, it is impossible to collect corresponding radar reflectivity in areas such as mountains and oceans. In such cases, using deep learning methods to reconstruct radar reflectivity from meteorological satellite data becomes an effective solution. However, Earth observation data is complex and exhibits strong long-range dependencies. Such data characteristics require the ability to model long-distance dependencies, making the Transformer more suitable for this specific data. With the ultimate goal of producing high reconstruction quality, we propose in this paper a novel architecture, designated by SU3plus, based on the Swin Transformer with a multi-scale feature fusion mechanism, while leveraging multiple channels of satellite data. Moreover, we define an appropriate loss function by combining the weighted versions of standard metrics, to take into account the data distribution imbalance and improve the reconstruction performance. Extensive experiments, carried out on the SEVIR storm dataset, confirm the effectiveness of the proposed approach compared to several state-of-the-art models. Zhixuan Zhou, Xintong Zhao, Mounir Kaaniche, Ding Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Diagnosability Verification and Enforcement in Labeled Petri Nets Under Sensor AttacksabstractThis article formalizes and solves the problems of diagnosability verification and enforcement in discrete event systems modeled with labeled Petri nets (LPNs) under sensor attacks. Given a plant, attackers work as a group in the framework of a coordinated distributed architecture and have the ability to edit some sensor readings to conceal the faults to confuse the operator. Furthermore, attackers necessarily remain furtive, i.e., their presence should not be discovered by the operator. In order to describe the set of all possible furtive attacks, a joint furtive diagnoser is established. We prove that an LPN under the above attacks is diagnosable if and only if its joint furtive diagnoser does not have the cycles composed of pairs of either faulty states and normal states, or faulty states and uncertain states. A new labeling function is proposed to enforce a plant to be diagnosable against as many attacks as possible. Examples are provided to illustrate the proposed method. Shaopeng Hu 0001, Zhiwu Li 0001, Ding Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolutionabstractEfficient transformer-based models have made remarkable progress in image super-resolution (SR). Most of these works mainly design elaborate structures to accelerate the inference of the transformer, where all feature tokens are propagated equally. However, they ignore the underlying characteristic of image content, i.e., various image regions have distinct restoration difficulties, especially for large images (2K-8K), failing to achieve adaptive inference. In this work, we propose an adaptive token sparsification transformer (AdaFormer) to speed up the model inference for image SR. Specifically, a texture-relevant sparse attention block with parallel global and local branches is introduced, aiming to integrate informative tokens from the global view instead of only in fixed local windows. Then, an early-exit strategy is designed to progressively halt tokens according to the token importance. To estimate the plausibility of each token, we adopt a lightweight confidence estimator, which is constrained by an uncertainty-guided loss to obtain a binary halting mask about the tokens. Experiments on large images have illustrated that our proposal reduces nearly 90% latency against SwinIR on Test8K, while maintaining a comparable performance. Xiaotong Luo, Zekun Ai, Qiuyuan Liang, Ding Liu 0001, Yuan Xie 0006, Yanyun Qu, Yun Fu 0001 |
AAAI | 4 |
| 2024 | Analysis of Bitcoin Fork by Colored Petri NetsabstractBitcoin is under the threat of fork since it operates with a distributed ledger. Predicting the fork probability in advance is beneficial for taking early action to avoid malicious attacks. In this study, we compose a colored Petri net model of Bitcoin. Our model consists of a given number of nodes, and each node has five subpages representing the node structure: proof of work, broadcast blocks, verify blocks, and the process of adding blocks to blockchain, respectively. Simulation results of fork probability can be easily obtained and analyzed by observing the data in the measuring components of subpages. The results show that our model correctly simulates the fork probability: on recent Bitcoin data, compared with the results of the wide-known SimBlock simulator, a difference of some 4.3% has been obtained. Thus, taking into account vivid graphical representation, our model has certain advantages for the developing techniques of attack avoidance. Ding Liu 0001, Tatiana R. Shmeleva, Dmitry Zaitsev 0001 |
SMC | 2 |
| 2023 | ShadowFormer: Global Context Helps Shadow RemovalabstractRecent deep learning methods have achieved promising results in image shadow removal. However, most of the existing approaches focus on working locally within shadow and non-shadow regions, resulting in severe artifacts around the shadow boundaries as well as inconsistent illumination between shadow and non-shadow regions. It is still challenging for the deep shadow removal model to exploit the global contextual correlation between shadow and non-shadow regions. In this work, we first propose a Retinex-based shadow model, from which we derive a novel transformer-based network, dubbed ShandowFormer, to exploit non-shadow regions to help shadow region restoration. A multi-scale channel attention framework is employed to hierarchically capture the global information. Based on that, we propose a Shadow-Interaction Module (SIM) with Shadow-Interaction Attention (SIA) in the bottleneck stage to effectively model the context correlation between shadow and non-shadow regions. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to evaluate the proposed method. Our method achieves state-of-the-art performance by using up to 150X fewer model parameters. Lanqing Guo, Siyu Huang, Ding Liu 0001, Hao Cheng 0016, Bihan Wen |
AAAI | 3 |
| 2023 | Boosting Video Super Resolution with Patch-Based Temporal Redundancy Optimization
Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Boyang Liang, Yu Guo 0006, Ding Liu 0001, Lean Fu, Fei Wang 0008 |
ICANN (7) | 7 |
| 2023 | Hierarchical Integration Diffusion Model for Realistic Image DeblurringabstractDiffusion models (DMs) have recently been introduced in image deblurring and exhibited promising performance, particularly in terms of details reconstruction. However, the diffusion model requires a large number of inference iterations to recover the clean image from pure Gaussian noise, which consumes massive computational resources. Moreover, the distribution synthesized by the diffusion model is often misaligned with the target results, leading to restrictions in distortion-based metrics. To address the above issues, we propose the Hierarchical Integration Diffusion Model (HI-Diff), for realistic image deblurring. Specifically, we perform the DM in a highly compacted latent space to generate the prior feature for the deblurring process. The deblurring process is implemented by a regression-based method to obtain better distortion accuracy. Meanwhile, the highly compact latent space ensures the efficiency of the DM. Furthermore, we design the hierarchical integration module to fuse the prior into the regression-based model from multiple scales, enabling better generalization in complex blurry scenarios. Comprehensive experiments on synthetic and real-world blur datasets demonstrate that our HI-Diff outperforms state-of-the-art methods. Code and trained models are available at https://github.com/zhengchen1999/HI-Diff. Zheng Chen 0014, Yulun Zhang 0001, Ding Liu 0001, Bin Xia 0014, Jinjin Gu, Linghe Kong, Xin Yuan 0002 |
NeurIPS | 3 |
| 2023 | Pyramid Attention Network for Image RestorationabstractAbstract Self-similarity refers to the image prior widely used in image restoration algorithms that small but similar patterns tend to occur at different locations and scales. However, recent advanced deep convolutional neural network-based methods for image restoration do not take full advantage of self-similarities by relying on self-attention neural modules that only process information at the same scale. To solve this problem, we present a novel Pyramid Attention module for image restoration, which captures long-range feature correspondences from a multi-scale feature pyramid. Inspired by the fact that corruptions, such as noise or compression artifacts, drop drastically at coarser image scales, our attention module is designed to be able to borrow clean signals from their “clean” correspondences at the coarser levels. The proposed pyramid attention module is a generic building block that can be flexibly integrated into various neural architectures. Its effectiveness is validated through extensive experiments on multiple image restoration tasks: image denoising, demosaicing, compression artifact reduction, and super resolution. Without any bells and whistles, our PANet (pyramid attention module with simple network backbones) can produce state-of-the-art results with superior accuracy and visual quality. Our code is available at https://github.com/SHI-Labs/Pyramid-Attention-Networks Yiqun Mei, Yuchen Fan 0001, Yulun Zhang 0001, Yuqian Zhou, Ding Liu 0001, Yun Fu 0001, Thomas S. Huang, Humphrey Shi |
Int. J. Comput. Vis. | 6 |
| 2022 | Unsupervised Low-light Image Enhancement with Decoupled NetworksabstractIn this paper, we tackle the problem of enhancing real-world low-light images with significant noise in an unsupervised fashion. Conventional unsupervised approaches focus primarily on illumination or contrast enhancement but fail to suppress the noise in real-world low-light images. To address this issue, we decouple this task into two sub-tasks: illumination enhancement and noise suppression. We propose a two-stage, fully unsupervised model to handle these tasks separately. In the noise suppression stage, we propose an illumination-aware denoising model so that real noise at different locations is removed with the guidance of the illumination conditions. To facilitate the unsupervised training, we construct pseudo triplet samples and propose an adaptive content loss correspondingly to preserve contextual details. To thoroughly evaluate the performance of the enhancement models, we build a new unpaired real-world low-light enhancement dataset. Extensive experiments show that our proposed method outperforms the state-of-the-art unsupervised methods concerning both illumination enhancement and noise reduction. Wei Xiong 0008, Ding Liu 0001, Xiaohui Shen, Jiebo Luo 0001 |
ICPR | 2 |
| 2022 | Adjustable Memory-efficient Image Super-resolution via Individual Kernel SparsityabstractThough single image super-resolution (SR) has witnessed incredible progress, the increasing model complexity impairs its applications in memory-limited devices. To solve this problem, prior arts have aimed to reduce the number of model parameters and sparsity has been exploited, which usually enforces the group sparsity constraint on the filter level and thus is not arbitrarily adjustable for satisfying the customized memory requirements. In this paper, we propose an individual kernel sparsity (IKS) method for memory-efficient and sparsity-adjustable image SR to aid deep network deployment in memory-limited devices. IKS performs model sparsity in the weight level that implicitly allocates the user-defined target sparsity to each individual kernel. To induce the kernel sparsity, a soft thresholding operation is used as a gating constraint for filtering the trivial weights. To achieve adjustable sparsity, a dynamic threshold learning algorithm is proposed, in which the threshold is updated by associated training with the network weight and is adaptively decayed with the guidance of the desired sparsity. This work essentially provides a dynamic parameter reassignment scheme with a given resource budget for an off-the-shelf SR model. Extensive experimental results demonstrate that IKS imparts considerable sparsity with negligible effect on SR quality. The code is available at: https://github.com/RaccoonDML/IKS. Xiaotong Luo, Mingliang Dai, Yulun Zhang 0001, Yuan Xie 0006, Ding Liu 0001, Yanyun Qu, Yun Fu 0001, Junping Zhang |
ACM Multimedia | 5 |
| 2022 | Adversarial Open Domain Adaptation for Sketch-to-Photo SynthesisabstractIn this paper, we explore open-domain sketch-to-photo translation, which aims to synthesize a realistic photo from a freehand sketch with its class label, even if the sketches of that class are missing in the training data. It is challenging due to the lack of training supervision and the large geometric distortion between the freehand sketch and photo domains. To synthesize the absent freehand sketches from photos, we propose a framework that jointly learns sketch-to-photo and photo-to-sketch generation. However, the generator trained from fake sketches might lead to unsatisfying results when dealing with sketches of missing classes, due to the domain gap between synthesized sketches and real ones. To alleviate this issue, we further propose a simple yet effective open-domain sampling and optimization strategy to "fool" the generator into treating fake sketches as real ones. Our method takes advantage of the learned sketch-to-photo and photo-to-sketch mapping of in-domain data and generalizes it to the open-domain classes. We validate our method on the Scribble and SketchyCOCO datasets. Compared with the recent competing methods, our approach shows impressive results in synthesizing realistic color, texture, and maintaining the geometric composition for various categories of open-domain sketches. Xiaoyu Xiang, Ding Liu 0001, Yiheng Zhu 0003, Xiaohui Shen, Jan P. Allebach |
WACV | 2 |
| 2021 | CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationabstractVideo instance segmentation is a complex task in which we need to detect, segment, and track each object for any given video. Previous approaches only utilize single-frame features for the detection, segmentation, and tracking of objects and they suffer in the video scenario due to several distinct challenges such as motion blur and drastic appearance change. To eliminate ambiguities introduced by only using single-frame features, we propose a novel comprehensive feature aggregation approach (CompFeat) to refine features atboth frame-level and object-level with temporal and spatial context information. The aggregation process is carefully designed with a new attention mechanism which significantly increases the discriminative power of the learned features. We further improve the tracking capability of our model through a siamese design by incorporating both feature similarities and spatial similarities. Experiments conducted on the YouTube-VIS dataset validate the effectiveness of proposed CompFeat. Ding Liu 0001, Thomas S. Huang, Humphrey Shi |
AAAI | 3 |
| 2021 | Progressive Temporal Feature Alignment Network for Video InpaintingabstractVideo inpainting aims to fill spatiotemporal "corrupted" regions with plausible content. To achieve this goal, it is necessary to find correspondences from neighbouring frames to faithfully hallucinate the unknown con-tent. Current methods achieve this goal through attention, flow-based warping, or 3D temporal convolution. However, flow-based warping can create artifacts when optical flow is not accurate, while temporal convolution may suffer from spatial misalignment. We propose ‘Progressive Temporal Feature Alignment Network’, which progressively enriches features extracted from the current frame with the feature warped from neighbouring frames using optical flow. Our approach corrects the spatial misalignment in the temporal feature propagation stage, greatly improving visual quality and temporal consistency of the inpainted videos. Using the proposed architecture, we achieve state-of-the-art performance on the DAVIS and FVI datasets compared to existing deep learning approaches. Code is available at https://github.com/MaureenZOU/TSAM. Xueyan Zou, Ding Liu 0001, Yong Jae Lee |
CVPR | 3 |
| 2021 | A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗abstractWe present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requirements, most existing approaches only cater to a specific task or use different architectures to address various tasks. Here we propose a unified framework based on Conditional Variational Auto-Encoder (CVAE), where we treat any arbitrary input as a masked motion series. Notably, by considering this problem as a conditional generation process, we estimate a parametric distribution of the missing regions based on the input conditions, from which to sample and synthesize the full motion series. To further allow the flexibility of manipulating the motion style of the generated series, we design an Action-Adaptive Modulation (AAM) to propagate the given semantic guidance through the whole sequence. We also introduce a cross-attention mechanism to exploit distant relations among decoder and encoder features for better realism and global consistency. We conducted extensive experiments on Human 3.6M and CMU-Mocap. The results show that our method produces coherent and realistic results for various motion synthesis tasks, with the synthesized motions distinctly adapted by the given action labels. Yujun Cai, Yiwei Wang 0001, Yiheng Zhu 0003, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Chuanxia Zheng, Sijie Yan, Henghui Ding, Xiaohui Shen, Ding Liu 0001, Nadia Magnenat-Thalmann |
ICCV | 12 |
| 2021 | Boosting Lightweight Single Image Super-resolution via Joint-distillationabstractThe rising of deep learning has facilitated the development of single image super-resolution (SISR). However, the growing burdensome model complexity and memory occupation severely hinder its practical deployments on resource-limited devices. In this paper, we propose a novel joint-distillation (JDSR) framework to boost the representation of various off-the-shelf lightweight SR models. The framework includes two stages: the superior LR generation and the joint-distillation learning. The superior LR is obtained from the HR image itself. With less than $300$K parameters, the peer network using superior LR as input can achieve comparable SR performance with large models, e.g., RCAN, with 15M parameters, which enables it as the input of peer network to save the training expense. The joint-distillation learning consists of internal self-distillation and external mutual learning. The internal self-distillation aims to achieve model self-boosting by transferring the knowledge from the deeper SR output to the shallower one. Specifically, each intermediate SR output is supervised by the HR image and the soft label from subsequent deeper outputs. To shrink the capacity gap between shallow and deep layers, a soft label generator is designed in a progressive backward fusion way with meta-learning for adaptive weight fine-tuning. The external mutual learning focuses on obtaining interaction information from a peer network in the process. Moreover, a curriculum learning strategy and a performance gap threshold are introduced for balancing the convergence rate of the original SR model and its peer network. Comprehensive experiments on benchmark datasets demonstrate that our proposal improves the performance of recent lightweight SR models by a large margin, with the same model architecture and inference expense. Xiaotong Luo, Qiuyuan Liang, Ding Liu 0001, Yanyun Qu |
ACM Multimedia | 3 |
| 2021 | EnlightenGAN: Deep Light Enhancement Without Paired SupervisionabstractDeep learning-based methods have achieved remarkable success in image restoration and enhancement, but are they still competitive when there is a lack of paired training data? As one such example, this paper explores the low-light image enhancement problem, where in practice it is extremely challenging to simultaneously take a low-light and a normal-light photo of the same visual scene. We propose a highly effective unsupervised generative adversarial network, dubbed EnlightenGAN, that can be trained without low/normal-light image pairs, yet proves to generalize very well on various real-world test images. Instead of supervising the learning using ground truth data, we propose to regularize the unpaired training using the information extracted from the input itself, and benchmark a series of innovations for the low-light image enhancement problem, including a global-local discriminator structure, a self-regularized perceptual loss fusion, and the attention mechanism. Through extensive experiments, our proposed approach outperforms recent methods under a variety of metrics in terms of visual quality and subjective user study. Thanks to the great flexibility brought by unpaired training, EnlightenGAN is demonstrated to be easily adaptable to enhancing real-world images from various domains. Our codes and pre-trained models are available at: https://github.com/VITA-Group/EnlightenGAN. Yifan Jiang 0001, Xinyu Gong, Ding Liu 0001, Yu Cheng 0001, Xiaohui Shen, Jianchao Yang, Pan Zhou 0001, Zhangyang Wang |
IEEE Trans. Image Process. | 3 |
| 2020 | Scale-Wise Convolution for Image RestorationabstractWhile scale-invariant modeling has substantially boosted the performance of visual recognition tasks, it remains largely under-explored in deep networks based image restoration. Naively applying those scale-invariant techniques (e.g., multi-scale testing, random-scale data augmentation) to image restoration tasks usually leads to inferior performance. In this paper, we show that properly modeling scale-invariance into neural networks can bring significant benefits to image restoration performance. Inspired from spatial-wise convolution for shift-invariance, “scale-wise convolution” is proposed to convolve across multiple scales for scale-invariance. In our scale-wise convolutional network (SCN), we first map the input image to the feature space and then build a feature pyramid representation via bi-linear down-scaling progressively. The feature pyramid is then passed to a residual network with scale-wise convolutions. The proposed scale-wise convolution learns to dynamically activate and aggregate features from different input scales in each residual building block, in order to exploit contextual information on multiple scales. In experiments, we compare the restoration accuracy and parameter efficiency among our model and many different variants of multi-scale neural networks. The proposed network with scale-wise convolution achieves superior performance in multiple image restoration tasks including image super-resolution, image denoising and image compression artifacts removal. Code and models are available at: https://github.com/ychfan/scn_sr. Yuchen Fan 0001, Ding Liu 0001, Thomas S. Huang |
AAAI | 3 |
| 2020 | Learning Progressive Joint Propagation for Human Motion Prediction
Yujun Cai, Lin Huang 0004, Yiwei Wang 0001, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Xu Yang 0021, Yiheng Zhu 0003, Xiaohui Shen, Ding Liu 0001, Jing Liu 0050, Nadia Magnenat-Thalmann |
ECCV (7) | 11 |
| 2020 | Neural Sparse Representation for Image RestorationabstractInspired by the robustness and efficiency of sparse representation in sparse coding based image restoration models, we investigate the sparsity of neurons in deep networks. Our method structurally enforces sparsity constraints upon hidden neurons. The sparsity constraints are favorable for gradient-based learning algorithms and attachable to convolution layers in various networks. Sparsity in neurons enables computation saving by only operating on non-zero components without hurting accuracy. Meanwhile, our method can magnify representation dimensionality and model capacity with negligible additional computation cost. Experiments show that sparse representation is crucial in deep neural networks for multiple image restoration tasks, including image super-resolution, image denoising, and image compression artifacts removal. Yuchen Fan 0001, Yiqun Mei, Yulun Zhang 0001, Yun Fu 0001, Ding Liu 0001, Thomas S. Huang |
NeurIPS | 6 |
| 2020 | DAVID: Dual-Attentional Video DeblurringabstractBlind video deblurring restores sharp frames from a blurry sequence without any prior. It is a challenging task because the blur due to camera shake, object movement and defocusing is heterogeneous in both temporal and spatial dimensions. Traditional methods train on datasets synthesized with a single level of blur, and thus do not generalize well across levels of blurriness. To address this challenge, we propose a dual attention mechanism to dynamically aggregate temporal cues for deblurring with an end-to-end trainable network structure. Specifically, an internal attention module adaptively selects the optimal temporal scales for restoring the sharp center frame. An external attention module adaptively aggregates and refines multiple sharp frame estimates, from several internal attention modules designed for different blur levels. To train and evaluate on more diverse blur severity levels, we propose a Challenging DVD dataset generated from the raw DVD video set by pooling frames with different temporal windows. Our framework achieves consistently better performance on this more challenging dataset while obtaining strongly competitive results on the original DVD benchmark. Extensive ablative studies and qualitative visualizations further demonstrate the advantage of our method in handling real video blur. Xiang Yu 0002, Ding Liu 0001, Manmohan Krishna Chandraker, Zhangyang Wang |
WACV | 3 |
| 2020 | Learning Simple Thresholded Features With Sparse Support RecoveryabstractDue to the shortcomings of the weakly supervised and fully supervised object detection (i.e., unsatisfactory performance and expensive annotations, respectively), leveraging partially labeled images in a cost-effective way to train an object detector has attracted much attention. In this paper, we formulate this challenging task as a missing bounding-boxes' object detection problem. Specifically, we develop a pseudo ground truth mining procedure to automatically find the missing bounding boxes for the unlabeled instances, called pseudo ground truths here, in the training data, and then combine the mined pseudo ground truths and the labeled annotations to train a fully supervised object detector. Furthermore, we propose an incremental learning framework to gradually incorporate the results of the trained fully supervised detector to improve the performance of the missing bounding-boxes' object detection. More importantly, we find an effective way to label the massive images with limited labors and funds, which is crucial when building a large-scale weakly/webly labeled dataset for object detection. The extensive experiments on the PASCAL VOC and COCO benchmarks demonstrate that our proposed method can narrow the gap between the fully supervised and weakly supervised object detectors, and outperform the previous state-of-the-art weakly supervised detectors by a large margin (more than 3% mAP absolutely) when the missing rate equals 0.9. Moreover, our proposed method with 30% missing bounding-box annotations can achieve comparable performance to some fully supervised detectors. Hongyu Xu, Zhangyang Wang, Haichuan Yang, Ding Liu 0001, Ji Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | FormNet: Formatted Learning for Image RestorationabstractIn this paper, we propose a deep CNN to tackle the image restoration problem by learning formatted information. Previous deep learning based methods directly learn the mapping from corrupted images to clean images, and may suffer from the gradient exploding/vanishing problems of deep neural networks. We propose to address the image restoration problem by learning the structured details and recovering the latent clean image together, from the shared information between the corrupted image and the latent image. In addition, instead of learning the pure difference (corruption), we propose to add a residual formatting layer and an adversarial block to format the information to structured one, which allows the network to converge faster and boosts the performance. Furthermore, we propose a cross-level loss net to ensure both pixel-level accuracy and semantic-level visual quality. Evaluations on public datasets show that the proposed method performs favorably against existing approaches quantitatively and qualitatively. Jianbo Jiao, Wei-Chih Tu, Ding Liu 0001, Shengfeng He, Rynson W. H. Lau, Thomas S. Huang |
IEEE Trans. Image Process. | 3 |
| 2020 | Connecting Image Denoising and High-Level Vision Tasks via Deep LearningabstractImage denoising and high-level vision tasks are usually handled independently in the conventional practice of computer vision, and their connection is fragile. In this paper, we cope with the two jointly and explore the mutual influence between them with the focus on two questions, namely (1) how image denoising can help improving high-level vision tasks, and (2) how the semantic information from high-level vision tasks can be used to guide image denoising. First for image denoising we propose a convolutional neural network in which convolutions are conducted in various spatial resolutions via downsampling and upsampling operations in order to fuse and exploit contextual information on different scales. Second we propose a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via backpropagation. We experimentally show that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network produces more visually appealing results. Extensive experiments demonstrate the benefit of exploiting image semantics simultaneously for image denoising and highlevel vision tasks via deep learning. The code is available online: https://github.com/Ding-Liu/DeepDenoising. Ding Liu 0001, Bihan Wen, Jianbo Jiao, Xianming Liu 0005, Zhangyang Wang, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2019 | Improving Object Detection from Scratch via Gated Feature Reuse
Humphrey Shi, NhatHai Phan, Rogério Feris, Liangliang Cao, Ding Liu 0001, Xinchao Wang, Thomas S. Huang, Marios Savvides |
BMVC | 7 |
| 2019 | Improving 3D Human Pose Estimation Via 3D Part Affinity Fieldsabstract3D human pose estimation from monocular images has become a heated area in computer vision recently. For years, most deep neural network based practices have adopted either an end-to-end approach, or a two-stage approach. An end-to-end network typically estimates 3D human poses directly from 2D input images, but it suffers from the shortage of 3D human pose data. It is also obscure to know if the inaccuracy stems from limited visual under-standing or 2D-to-3D mapping. Whereas a two-stage directly lifts those 2D keypoint outputs to the 3D space, after utilizing an existing network for 2D keypoint detections. However, they tend to ignore some useful contextual hints from the 2D raw image pixels. In this paper, we introduce a two-stage architecture that can eliminate the main disadvantages of both these approaches. During the first stage we use an existing state-of-the-art detector to estimate 2D poses. To add more con-textual information to help lifting 2D poses to 3D poses, we propose 3D Part Affinity Fields (3D-PAFs). We use 3D-PAFs to infer 3D limb vectors, and combine them with 2D poses to regress the 3D coordinates. We trained and tested our proposed framework on Human3.6M, the most popular 3D human pose benchmark dataset. Our approach achieves the state-of-the-art performance, which proves that with right selections of contextual information, a simple regression model can be very powerful in estimating 3D poses. Ding Liu 0001, Xinchao Wang, Yuxiao Hu 0001, Lei Zhang 0001, Thomas S. Huang |
WACV | 1 |
| 2019 | Enhance Visual Recognition Under Adverse Conditions via Deep NetworksabstractVisual recognition under adverse conditions is a very important and challenging problem of high practical value, due to the ubiquitous existence of quality distortions during image acquisition, transmission, or storage. While deep neural networks have been extensively exploited in the techniques of low-quality image restoration and high-quality image recognition tasks respectively, few studies have been done on the important problem of recognition from very low-quality images. This paper proposes a deep learning based framework for improving the performance of image and video recognition models under adverse conditions, using robust adverse pre-training or its aggressive variant. The robust adverse pre-training algorithms leverage the power of pre-training and generalizes conventional unsupervised pre-training and data augmentation methods. We further develop a transfer learning approach to cope with real-world datasets of unknown adverse conditions. The proposed framework is comprehensively evaluated on a number of image and video recognition benchmarks, and obtains significant performance improvements under various single or mixed adverse conditions. Our visualization and analysis further add to the explainability of results. Ding Liu 0001, Bowen Cheng, Zhangyang Wang, Haichao Zhang 0001, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2018 | Visual Recognition in Very Low-Quality Settings: Delving Into the Power of Pre-TrainingabstractVisual recognition from very low-quality images is an extremely challenging task with great practical values. While deep networks have been extensively applied to low-quality image restoration and high-quality image recognition tasks respectively, few works have been done on the important problem of recognition from very low-quality images.This paper presents a degradation-robust pre-training approach on improving deep learning models towards this direction. Extensive experiments on different datasets validate the effectiveness of our proposed method. Bowen Cheng, Ding Liu 0001, Zhangyang Wang, Haichao Zhang 0001, Thomas S. Huang |
AAAI | 2 |
| 2018 | Image Super-Resolution via Dual-State Recurrent NetworksabstractAdvances in image super-resolution (SR) have recently benefited significantly from rapid developments in deep neural networks. Inspired by these recent discoveries, we note that many state-of-the-art deep SR architectures can be reformulated as a single-state recurrent neural network (RNN) with finite unfoldings. In this paper, we explore new structures for SR based on this compact RNN view, leading us to a dual-state design, the Dual-State Recurrent Network (DSRN). Compared to its single-state counterparts that operate at a fixed spatial resolution, DSRN exploits both low-resolution (LR) and high-resolution (HR) signals jointly. Recurrent signals are exchanged between these states in both directions (both LR to HR and HR to LR) via delayed feedback. Extensive quantitative and qualitative evaluations on benchmark datasets and on a recent challenge demonstrate that the proposed DSRN performs favorably against state-of-the-art algorithms in terms of both memory consumption and predictive accuracy. The code for our method is publicly available1. Wei Han 0002, Shiyu Chang, Ding Liu 0001, Mo Yu, Michael Witbrock, Thomas S. Huang |
CVPR | 3 |
| 2018 | Survey of Face Detection on Low-Quality ImagesabstractFace detection is a well-explored problem. Many challenges on face detectors like extreme pose, illumination, low resolution and small scales are studied in the previous work. However, previous proposed models are mostly trained and tested on good-quality images which are not always the case for practical applications like surveillance systems. In this paper, we first review the current state-of-the-art face detectors and their performance on benchmark dataset FDDB, and compare the design protocols of the algorithms. Secondly, we investigate their performance degradation while testing on low-quality images with different levels of blur, noise, and contrast. Our results demonstrate that both hand-crafted and deep-learning based face detectors are not robust enough for low-quality images. It inspires researchers to produce more robust design for face detection in the wild. Yuqian Zhou, Ding Liu 0001, Thomas S. Huang |
FG | 2 |
| 2018 | When Image Denoising Meets High-Level Vision Tasks: A Deep Learning ApproachabstractConventionally, image denoising and high-level vision tasks are handled separately in computer vision. In this paper, we cope with the two jointly and explore the mutual influence between them. First we propose a convolutional neural network for image denoising which achieves the state-of-the-art performance. Second we propose a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via back-propagation. We demonstrate that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network can generate more visually appealing results. To the best of our knowledge, this is the first work investigating the benefit of exploiting image semantics simultaneously for image denoising and high-level vision tasks via deep learning. Ding Liu 0001, Bihan Wen, Xianming Liu 0005, Zhangyang Wang, Thomas S. Huang |
IJCAI | 1 |
| 2018 | Non-Local Recurrent Network for Image RestorationabstractMany classic methods have shown non-local self-similarity in natural images to be an effective prior for image restoration. However, it remains unclear and challenging to make use of this intrinsic property via deep networks. In this paper, we propose a non-local recurrent network (NLRN) as the first attempt to incorporate non-local operations into a recurrent neural network (RNN) for image restoration. The main contributions of this work are: (1) Unlike existing methods that measure self-similarity in an isolated manner, the proposed non-local module can be flexibly integrated into existing deep networks for end-to-end training to capture deep feature correlation between each location and its neighborhood. (2) We fully employ the RNN structure for its parameter efficiency and allow deep feature correlation to be propagated along adjacent recurrent states. This new design boosts robustness against inaccurate correlation estimation due to severely degraded images. (3) We show that it is essential to maintain a confined neighborhood for computing deep feature correlation given degraded images. This is in contrast to existing practice that deploys the whole image. Extensive experiments on both image denoising and super-resolution tasks are conducted. Thanks to the recurrent non-local operations and correlation propagation, the proposed NLRN achieves superior results to state-of-the-art methods with many fewer parameters. Ding Liu 0001, Bihan Wen, Yuchen Fan 0001, Chen Change Loy, Thomas S. Huang |
NeurIPS | 1 |
| 2018 | Understanding Convolution for Semantic SegmentationabstractRecent advances in deep learning, especially deep convolutional neural networks (CNNs), have led to significant improvement over previous semantic segmentation systems. Here we show how to improve pixel-wise semantic segmentation by manipulating convolution-related operations that are of both theoretical and practical value. First, we design dense upsampling convolution (DUC) to generate pixel-level prediction, which is able to capture and decode more detailed information that is generally missing in bilinear upsampling. Second, we propose a hybrid dilated convolution (HDC) framework in the encoding phase. This framework 1) effectively enlarges the receptive fields (RF) of the network to aggregate global information; 2) alleviates what we call the "gridding issue"caused by the standard dilated convolution operation. We evaluate our approaches thoroughly on the Cityscapes dataset, and achieve a state-of-art result of 80.1% mIOU in the test set at the time of submission. We also have achieved state-of-theart overall on the KITTI road estimation benchmark and the PASCAL VOC2012 segmentation task. Our source code can be found at https://github.com/TuSimple/TuSimple-DUC. Panqu Wang, Ding Liu 0001, Zehua Huang, Garrison W. Cottrell |
WACV | 4 |
| 2018 | Learning Temporal Dynamics for Video Super-Resolution: A Deep Learning ApproachabstractVideo super-resolution (SR) aims at estimating a high-resolution (HR) video sequence from a low-resolution (LR) one. Given that deep learning has been successfully applied to the task of single image SR, which demonstrates the strong capability of neural networks for modeling spatial relation within one single image, the key challenge to conduct video SR is how to efficiently and effectively exploit the temporal dependency among consecutive LR frames other than the spatial relation. However, this remains challenging because complex motion is difficult to model and can bring detrimental effects if not handled properly. We tackle the problem of learning temporal dynamics from two aspects. First, we propose a temporal adaptive neural network that can adaptively determine the optimal scale of temporal dependency. Inspired by the Inception module in GoogLeNet [1], filters of various temporal scales are applied to the input LR sequence before their responses are adaptively aggregated, in order to fully exploit the temporal relation among consecutive LR frames. Second, we decrease the complexity of motion among neighboring frames using a spatial alignment network that can be end-to-end trained with the temporal adaptive network and has the merit of increasing the robustness to complex motion and the efficiency compared to competing image alignment methods. We provide a comprehensive evaluation of the temporal adaptation and the spatial alignment modules. We show the temporal adaptive design considerably improve SR quality over its plain counterparts, and the spatial alignment network is able to attain comparable SR performance with the sophisticated optical flow based approach, but requires much less running time. Overall our proposed model with learned temporal dynamics is shown to achieve state-of-the-art SR results in terms of not only spatial consistency but also temporal coherence on public video datasets. More information can be found in. Ding Liu 0001, Yuchen Fan 0001, Xianming Liu 0005, Zhangyang Wang, Shiyu Chang, Xinchao Wang, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2017 | Robust emotion recognition from low quality and low bit rate video: A deep learning approachabstractEmotion recognition from facial expressions is tremendously useful, especially when coupled with smart devices and wireless multimedia applications. However, the inadequate network bandwidth often limits the spatial resolution of the transmitted video, which will heavily degrade the recognition reliability. We develop a novel framework to achieve robust emotion recognition from low bit rate video. While video frames are downsampled at the encoder side, the decoder is embedded with a deep network model for joint super-resolution (SR) and recognition. Notably, we propose a novel max-mix training strategy, leading to a single “One-for-All” model that is remarkably robust to a vast range of downsampling factors. That makes our framework well adapted for the varied bandwidths in real transmission scenarios, without hampering scalability or efficiency. The proposed framework is evaluated on the AVEC 2016 benchmark, and demonstrates significantly improved stand-alone recognition performance, as well as rate-distortion (R-D) performance, than either directly recognizing from LR frames, or separating SR and recognition. Bowen Cheng, Zhangyang Wang, Zhaobin Zhang, Zhu Li 0001, Ding Liu 0001, Jianchao Yang, Shuai Huang 0001, Thomas S. Huang |
ACII | 5 |
| 2017 | Robust Video Super-Resolution with Learned Temporal DynamicsabstractVideo super-resolution (SR) aims to generate a high-resolution (HR) frame from multiple low-resolution (LR) frames in a local temporal window. The inter-frame temporal relation is as crucial as the intra-frame spatial relation for tackling this problem. However, how to utilize temporal information efficiently and effectively remains challenging since complex motion is difficult to model and can introduce adverse effects if not handled properly. We address this problem from two aspects. First, we propose a temporal adaptive neural network that can adaptively determine the optimal scale of temporal dependency. Filters on various temporal scales are applied to the input LR sequence before their responses are adaptively aggregated. Second, we reduce the complexity of motion between neighboring frames using a spatial alignment network which is much more robust and efficient than competing alignment methods and can be jointly trained with the temporal adaptive network in an end-to-end manner. Our proposed models with learned temporal dynamics are systematically evaluated on public video datasets and achieve state-of-the-art SR results compared with other recent video SR approaches. Both of the temporal adaptation and the spatial alignment modules are demonstrated to considerably improve SR quality over their plain counterparts. Ding Liu 0001, Yuchen Fan 0001, Xianming Liu 0005, Zhangyang Wang, Shiyu Chang, Thomas S. Huang |
ICCV | 1 |
| 2017 | Computed tomography super-resolution using convolutional neural networksabstractThe practical application of Computed Tomography (CT) faces the dilemma between higher image resolution and less X-ray exposure for patients, motivating the research on CT super-resolution (SR). In this paper, we apply state-of-the-art SR techniques to reconstruct CT images using two proposed advanced CT SR models based on Convolutional Neural Networks (CNNs) and residual learning: a single-slice CT SR network (S-CTSRN), and a multi-slice CT SR network (M-CTSRN). S-CTSRN improves the high-frequency feature extraction by incorporating the residual learning strategy, while M-CTSRN further utilizes the coherence between neighboring CT slices for better SR reconstruction. We evaluate both models on a large-scale CT dataset1, and obtain competitive results both quantitatively and qualitatively. Haichao Yu, Ding Liu 0001, Humphrey Shi, Hanchao Yu, Zhangyang Wang, Xinchao Wang, Brent Cross, Matthew Bramler, Thomas S. Huang |
ICIP | 2 |
| 2017 | Image aesthetics assessment using Deep Chatterjee's machineabstractImage aesthetics assessment has been challenging due to its subjective nature. Inspired by the Chatterjee's visual neuroscience model, we design Deep Chatterjee's Machine (DCM) tailored for this task. DCM first learns attributes through the parallel supervised pathways, on a variety of selected feature dimensions. A high-level synthesis network is trained to associate and transform those attributes into the overall aesthetics rating. We then extend DCM to predicting the distribution of human ratings, since aesthetics ratings are often subjective. We also highlight our first-of-its-kind study of label-preserving transformations in the context of aesthetics assessment, which leads to an effective data augmentation approach. Experimental results on the AVA dataset show that DCM gains significant performance improvement, compared to other state-of-the-art models. Zhangyang Wang, Ding Liu 0001, Shiyu Chang, Florin Dolcos, Diane M. Beck, Thomas S. Huang |
IJCNN | 2 |
| 2017 | Deadlock and liveness characterization for a class of generalized Petri nets
ShouGuang Wang, MengChu Zhou, Ding Liu 0001, Abdulrahman Al-Ahmari, Ting Qu 0002, Zhiwu Li 0001 |
Inf. Sci. | 4 |
| 2016 | Epitomic Image Super-ResolutionabstractWe propose Epitomic Image Super-Resolution (ESR) to enhance the current internal SR methods that exploit the self-similarities in the input. Instead of local nearest neighbor patch matching used in most existing internal SR methods, ESR employs epitomic patch matching that features robustness to noise, and both local and non-local patch matching. Extensive objective and subjective evaluation demonstrate the effectiveness and advantage of ESR on various images. Yingzhen Yang, Zhangyang Wang, Shiyu Chang, Ding Liu 0001, Humphrey Shi, Thomas S. Huang |
AAAI | 5 |
| 2016 | Learning a Mixture of Deep Networks for Single Image Super-Resolution
Ding Liu 0001, Nasser M. Nasrabadi, Thomas S. Huang |
ACCV (3) | 1 |
| 2016 | Studying Very Low Resolution Recognition Using Deep NetworksabstractVisual recognition research often assumes a sufficient resolution of the region of interest (ROI). That is usually violated in practice, inspiring us to explore the Very Low Resolution Recognition (VLRR) problem. Typically, the ROI in a VLRR problem can be smaller than 16 16 pixels, and is challenging to be recognized even by human experts. We attempt to solve the VLRR problem using deep learning methods. Taking advantage of techniques primarily in super resolution, domain adaptation and robust regression, we formulate a dedicated deep learning method and demonstrate how these techniques are incorporated step by step. Any extra complexity, when introduced, is fully justified by both analysis and simulation results. The resulting Robust Partially Coupled Networks achieves feature enhancement and recognition simultaneously. It allows for both the flexibility to combat the LR-HR domain mismatch, and the robustness to outliers. Finally, the effectiveness of the proposed models is evaluated on three different VLRR tasks, including face identification, digit recognition and font recognition, all of which obtain very impressive performances. Zhangyang Wang, Shiyu Chang, Yingzhen Yang, Ding Liu 0001, Thomas S. Huang |
CVPR | 4 |
| 2016 | D3: Deep Dual-Domain Based Fast Restoration of JPEG-Compressed ImagesabstractIn this paper, we design a Deep Dual-Domain (D3) based fast restoration model to remove artifacts of JPEG compressed images. It leverages the large learning capacity of deep networks, as well as the problem-specific expertise that was hardly incorporated in the past design of deep architectures. For the latter, we take into consideration both the prior knowledge of the JPEG compression scheme, and the successful practice of the sparsity-based dual-domain approach. We further design the One-Step Sparse Inference (1-SI) module, as an efficient and lightweighted feed-forward approximation of sparse coding. Extensive experiments verify the superiority of the proposed D3 model over several state-of-the-art methods. Specifically, our best model is capable of outperforming the latest deep model for around 1 dB in PSNR, and is 30 times faster. Zhangyang Wang, Ding Liu 0001, Shiyu Chang, Qing Ling 0001, Yingzhen Yang, Thomas S. Huang |
CVPR | 2 |
| 2016 | Robust Single Image Super-Resolution via Deep Networks With Sparse PriorabstractSingle image super-resolution (SR) is an ill-posed problem, which tries to recover a high-resolution image from its low-resolution observation. To regularize the solution of the problem, previous methods have focused on designing good priors for natural images, such as sparse representation, or directly learning the priors from a large data set with models, such as deep neural networks. In this paper, we argue that domain expertise from the conventional sparse coding model can be combined with the key ingredients of deep learning to achieve further improved results. We demonstrate that a sparse coding model particularly designed for SR can be incarnated as a neural network with the merit of end-to-end optimization over training data. The network has a cascaded structure, which boosts the SR performance for both fixed and incremental scaling factors. The proposed training and testing schemes can be extended for robust handling of images with additional degradation, such as noise and blurring. A subjective assessment is conducted and analyzed in order to thoroughly evaluate various SR techniques. Our proposed model is tested on a wide range of images, and it significantly outperforms the existing state-of-the-art methods for various scaling factors both quantitatively and perceptually. Ding Liu 0001, Bihan Wen, Jianchao Yang, Wei Han 0002, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2015 | Deep Networks for Image Super-Resolution with Sparse PriorabstractDeep learning techniques have been successfully applied in many areas of computer vision, including low-level image restoration problems. For image super-resolution, several models based on deep neural networks have been recently proposed and attained superior performance that overshadows all previous handcrafted models. The question then arises whether large-capacity and data-driven models have become the dominant solution to the ill-posed super-resolution problem. In this paper, we argue that domain expertise represented by the conventional sparse coding model is still valuable, and it can be combined with the key ingredients of deep learning to achieve further improved results. We show that a sparse coding model particularly designed for super-resolution can be incarnated as a neural network, and trained in a cascaded structure from end to end. The interpretation of the network based on sparse coding leads to much more efficient and effective training, as well as a reduced model size. Our model is evaluated on a wide range of images, and shows clear advantage over existing state-of-the-art methods in terms of both restoration accuracy and human subjective quality. Ding Liu 0001, Jianchao Yang, Wei Han 0002, Thomas S. Huang |
ICCV | 2 |
| 2013 | Hybrid Liveness-Enforcing Policy for Generalized Petri Net Models of Flexible Manufacturing SystemsabstractThis paper proposes a hybrid liveness-enforcing method for a class of Petri nets, which can well model many flexible manufacturing systems. The proposed method combines elementary siphons with a characteristic structure-based method to prevent deadlocks and enforce liveness to the net class under consideration. The characteristic structure-based method is further advanced in this work. It unveils and takes a full advantage of an intrinsically live structure of generalized Petri nets, which hides behind the arc weights, to achieve the liveness enforcement without any external control agent such as monitors. This hybrid method can identify and remove redundant monitors from a liveness-enforcing supervisor designed according to existing policies, improve the permissiveness, reduce the structural complexity of a controlled system, and consequently save the control implementation cost. Several examples are used to illustrate this method. Ding Liu 0001, Zhiwu Li 0001, MengChu Zhou |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |