EDBT 2026 Demo / reviewers in the wild / expert
Youliang Yan
dblp:135/5316
· DBLP profile ↗
34ranked-venue papers
0as first author
23since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 17 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Spectral Image Color ReproductionabstractFrom camera to screen, researchers have developed a well-established system for capturing and reproducing the color experience of human eyes. In this study, we aim to upgrade this process by transiting from conventional RGB to multi-spectral image (MSI) color reproduction. While MSI offers evident advantages in color matching, we find out it is not trivial to make good use of more spectral information for color constancy. Therefore, we present a regularized color reproduction system that incorporates a spectral prior-guided optimization strategy to establish a sensor-optimized RGB projection for color matching, along with a learning-based chromatic adaptation model for color constancy. Specifically, we define the RGB projection through an end-to-end optimization under the guidance of sensor spectral sensitivities. Subsequently, we devise a chromatic adaptation neural network that estimates the scene illuminance and an illuminance-adaptive matrix for auto white balancing and dynamic color correction, respectively. Comprehensive experiments show the superiority of our system compared to alternative solutions. Jiacheng Li 0004, Chang Chen 0004, Fenglong Song, Youliang Yan, Zhiwei Xiong |
WACV | 5 |
| 2025 | LeRF: Learning Resampling Function for Adaptive and Efficient Image InterpolationabstractImage resampling is a basic technique that is widely employed in daily applications, such as camera photo editing. Recent deep neural networks (DNNs) have made impressive progress in performance by introducing learned data priors. Still, these methods are not the perfect substitute for interpolation, due to the drawbacks in efficiency and versatility. In this work, we propose a novel method of Learning Resampling Function (termed LeRF), which takes advantage of both the structural priors learned by DNNs and the locally continuous assumption of interpolation. Specifically, LeRF assigns spatially varying resampling functions to input image pixels and learns to predict the hyper-parameters that determine the shapes of these resampling functions with a neural network. Based on the formulation of LeRF, we develop a family of models, including both efficiency-orientated and performance-orientated ones. To achieve interpolation-level efficiency, we adopt look-up tables (LUTs) to accelerate the inference of the learned neural network. Furthermore, we design a directional ensemble strategy and edge-sensitive indexing patterns to better capture local structures. On the other hand, to obtain DNN-level performance, we propose an extension of LeRF to enable it in cooperation with pre-trained upsampling models for cascaded resampling. Extensive experiments show that the efficiency-orientated version of LeRF runs as fast as interpolation, generalizes well to arbitrary transformations, and outperforms interpolation significantly, e.g., up to 3 dB PSNR gain over Bicubic for $\times 2$×2 upsampling on Manga109. Besides, the performance-orientated version of LeRF reaches comparable performance with existing DNNs at much higher efficiency, e.g., less than 25% running time on a desktop GPU. Jiacheng Li 0004, Chang Chen 0004, Fenglong Song, Youliang Yan, Zhiwei Xiong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised LearningabstractFor image super-resolution (SR), bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel “Low-Res Leads the Way” (LWay) training framework, merging Supervised Pre-training with Self-supervised Learning to enhance the adaptability of SR models to real-world images. Our approach utilizes a low-resolution (LR) reconstruction network to extract degradation embeddings from LR images, merging them with super-resolved outputs for LR reconstruction. Leveraging unseen LR images for self-supervised learning guides the model to adapt its modeling space to the target domain, facili-tating fine-tuning of SR models without requiring paired high-resolution (HR) images. The integration of Discrete Wavelet Transform (DWT)further refines the focus on high-frequency details. Extensive evaluations show that our method significantly improves the generalization and de-tail restoration capabilities of SR models on unseen real-world datasets, outperforming existing methods. Our training regime is universally compatible, requiring no network architecture modifications, making it a practical solution for real-world SR applications. Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Haoze Sun, Xueyi Zou, Zhensong Zhang, Youliang Yan, Lei Zhu 0003 |
CVPR | 8 |
| 2024 | VastGaussian: Vast 3D Gaussians for Large Scene ReconstructionabstractExisting NeRF-based methods for large scene reconstruction often have limitations in visual quality and rendering speed. While the recent 3D Gaussian Splatting works well on small-scale and object-centric scenes, scaling it up to large scenes poses challenges due to limited video memory, long optimization time, and noticeable appearance variations. To address these challenges, we present VastGaussian, the first method for high-quality reconstruction and real-time rendering on large scenes based on 3D Gaussian Splatting. We propose a progressive partitioning strategy to divide a large scene into multiple cells, where the training cameras and point cloud are properly distributed with an airspace-aware visibility criterion. These cells are merged into a complete scene after parallel optimization. We also introduce decoupled appearance modeling into the optimization process to reduce appearance variations in the rendered images. Our approach outperforms existing NeRF-based methods and achieves state-of-the-art results on multiple large scene datasets, enabling fast optimization and high-fidelity real-time rendering. Project page: https://vastgaussian.github.io. Zhihao Li 0002, Jianzhuang Liu, Shiyong Liu, Jiayue Liu, Yangdi Lu, Songcen Xu, Youliang Yan, Wenming Yang |
CVPR | 10 |
| 2024 | CoSeR: Bridging Image and Language for Cognitive Super-ResolutionabstractExisting super-resolution (SR) models primarily focus on restoring local texture details, often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the intro-duction of inaccurate textures during the recovery process. In our work, we introduce the Cognitive Super-Resolution (CoSeR) framework, empowering SR models with the ca-pacity to comprehend low-resolution images. We achieve this by marrying image appearance and language under-standing to generate a cognitive embedding, which not only activates prior information from large text-to-image diffusion models but also facilitates the generation of high-quality reference images to optimize the SR process. To fur-ther improve image fidelity, we propose a novel condition injection scheme called “Ali-in-Attention ”, consolidating all conditional information into a single module. Conse-quently, our method successfully restores semantically cor-rect and photorealistic details, demonstrating state-of-the-art performance across multiple benchmarks. Project page: https://coser-main.github.io/ Haoze Sun, Wenbo Li 0002, Jianzhuang Liu, Haoyu Chen 0003, Renjing Pei, Xueyi Zou, Youliang Yan, Yujiu Yang 0001 |
CVPR | 7 |
| 2024 | GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction
Yuxuan Mu, Xinxin Zuo, Chuan Guo 0002, Juwei Lu, Songcen Xu, Peng Dai 0002, Youliang Yan, Li Cheng 0001 |
ECCV (79) | 9 |
| 2024 | Learning Exhaustive Correlation for Spectral Super-Resolution: Where Spatial-Spectral Attention Meets Linear Dependence
Lizhi Wang 0001, Chang Chen 0004, Fenglong Song, Youliang Yan |
ECCV (25) | 7 |
| 2024 | Generative Human Motion Stylization in Latent SpaceabstractHuman motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction and infusion. Building upon this, we present a novel \textit{generative} model that produces diverse stylization results of a single motion (latent) code. During training, a motion code is decomposed into two coding components: a deterministic content code, and a probabilistic style code adhering to a prior distribution; then a generator massages the random combination of content and style codes to reconstruct the corresponding motion codes. Our approach is versatile, allowing the learning of probabilistic style space from either style labeled or unlabeled motions, providing notable flexibility in stylization as well. In inference, users can opt to stylize a motion using style cues from a reference motion or a label. Even in the absence of explicit style input, our model facilitates novel re-stylization by sampling from the unconditional style prior distribution. Experimental results show that our proposed stylization models, despite their lightweight design, outperform the state-of-the-arts in style reeanactment, content preservation, and generalization across various applications and settings. Chuan Guo 0002, Yuxuan Mu, Xinxin Zuo, Peng Dai 0002, Youliang Yan, Juwei Lu, Li Cheng 0001 |
ICLR | 5 |
| 2024 | Wavelet-based network for high dynamic range imagingabstractHigh dynamic range (HDR) imaging from multiple low dynamic range (LDR) images has been suffering from ghosting artifacts caused by scene and objects motion. Existing methods, such as optical flow based and end-to-end deep learning based solutions, are error-prone either in detail restoration or ghosting artifacts removal. Comprehensive empirical evidence shows that ghosting artifacts caused by large foreground motion are mainly low-frequency signals and the details are mainly high-frequency signals. In this work, we propose a novel frequency-guided end-to-end deep neural network (FHDRNet) to conduct HDR fusion in the frequency domain, and Discrete Wavelet Transform (DWT) is used to decompose inputs into different frequency bands. The low-frequency signals are used to avoid specific ghosting artifacts, while the high-frequency signals are used for preserving details. Using a U-Net as the backbone, we propose two novel modules: merging module and frequency-guided upsampling module. The merging module applies the attention mechanism to the low-frequency components to deal with the ghost caused by large foreground motion. The frequency-guided upsampling module reconstructs details from multiple frequency-specific components with rich details. In addition, a new RAW dataset is created for training and evaluating multi-frame HDR imaging algorithms in the RAW domain. Extensive experiments are conducted on public datasets and our RAW dataset, showing that the proposed FHDRNet achieves state-of-the-art performance. Tianhong Dai, Wei Li 0002, Xilei Cao, Jianzhuang Liu, Xu Jia 0012, Ales Leonardis, Youliang Yan, Shanxin Yuan |
Comput. Vis. Image Underst. | 7 |
| 2023 | Style Projected Clustering for Domain Generalized Semantic SegmentationabstractExisting semantic segmentation methods improve generalization capability, by regularizing various images to a canonical feature space. While this process contributes to generalization, it weakens the representation inevitably. In contrast to existing methods, we instead utilize the difference between images to build a better representation space, where the distinct style features are extracted and stored as the bases of representation. Then, the generalization to unseen image styles is achieved by projecting features to this known space. Specifically, we realize the style projection as a weighted combination of stored bases, where the similarity distances are adopted as the weighting factors. Based on the same concept, we extend this process to the decision part of model and promote the generalization of semantic prediction. By measuring the similarity distances to semantic bases (i.e., prototypes), we replace the common deterministic prediction with semantic clustering. Comprehensive experiments demonstrate the advantage of proposed method to the state of the art, up to 3.6% mIoU improvement in average on unseen scenarios. Code and models are available at https://gitee.com/mindspore/models/tree/master/research/cv/SPC-Net. Wei Huang 0036, Chang Chen 0004, Jiacheng Li 0004, Cheng Li 0009, Fenglong Song, Youliang Yan, Zhiwei Xiong |
CVPR | 7 |
| 2023 | Learning Steerable Function for Efficient Image ResamplingabstractImage resampling is a basic technique that is widely employed in daily applications. Existing deep neural networks (DNNs) have made impressive progress in resampling performance. Yet these methods are still not the perfect substitute for interpolation, due to the issues of efficiency and continuous resampling. In this work, we propose a novel method of Learning Resampling Function (termed LeRF), which takes advantage of both the structural priors learned by DNNs and the locally continuous assumption of interpolation methods. Specifically, LeRF assigns spatially-varying steerable resampling functions to input image pixels and learns to predict the hyper-parameters that determine the orientations of these resampling functions with a neural network. To achieve highly efficient inference, we adopt look-up tables (LUTs) to accelerate the inference of the learned neural network. Furthermore, we design a directional ensemble strategy and edge-sensitive indexing patterns to better capture local structures. Extensive experiments show that our method runs as fast as interpolation, generalizes well to arbitrary transformations, and outperforms interpolation significantly, e.g., up to 3dB PSNR gain over bicubic for x 2 upsampling on Manga109. Jiacheng Li 0004, Chang Chen 0004, Wei Huang 0036, Zhiqiang Lang, Fenglong Song, Youliang Yan, Zhiwei Xiong |
CVPR | 6 |
| 2023 | CLIPPING: Distilling CLIP-Based Models with a Student Base for Video-Language RetrievalabstractPre-training a vision-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vision-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capability of a large model to a small one while maintaining the accuracy, which has achieved remarkable success in natural language processing. However, it faces many problems when applying KD to the multi-modality applications. In this paper, we propose a novel knowledge distillation method, named CLIPPING11In this paper, CLIPPING means cutting something to make it smaller through distilling., where the plentiful knowledge of a large teacher model that has been fine-tuned for video-language tasks with the powerful pre-trained CLIP can be effectively transferred to a small student only at the fine-tuning stage. Especially, a new layer-wise alignment with the student as the base is proposed for knowledge distillation of the intermediate layers in CLIPPING, which enables the student's layers to be the bases of the teacher, and thus allows the student to fully absorb the knowledge of the teacher. CLIPPING with MobileViT-v2 as the vision encoder without any vision-language pre-training achieves 88.1%-95.3% of the performance of its teacher on three video-language retrieval benchmarks, with its vision encoder being 19.5x smaller. CLIPPING also significantly outperforms a state-of-the-art small baseline (ALL-in-one-B) on the MSR-VTT dataset, obtaining relatively 7.4% performance gain, with 29% fewer parameters and 86.9% fewer flops. Moreover, CLIPPING is comparable or even superior to many large pre-training models. Renjing Pei, Jianzhuang Liu, Weimian Li, Songcen Xu, Peng Dai 0002, Juwei Lu, Youliang Yan |
CVPR | 8 |
| 2023 | Toward RAW Object Detection: A New Benchmark and A New ModelabstractIn many computer vision applications (e.g., robotics and autonomous driving), high dynamic range (HDR) data is necessary for object detection algorithms to handle a variety of lighting conditions, such as strong glare. In this paper, we aim to achieve object detection on RAW sensor data, which naturally saves the HDR information from image sensors without extra equipment costs. We build a novel RAW sensor dataset, named ROD, for Deep Neural Networks (DNNs)-based object detection algorithms to be applied to HDR data. The ROD dataset contains a large amount of annotated instances of day and night driving scenes in 24-bit dynamic range. Based on the dataset, we first investigate the impact of dynamic range for DNNs-based detectors and demonstrate the importance of dynamic range adjustment for detection on RAW sensor data. Then, we propose a simple and effective adjustment method for object detection on HDR RAW sensor data, which is image adaptive and jointly optimized with the downstream detector in an end-to-end scheme. Extensive experiments demonstrate that the performance of detection on RAW sensor data is significantly superior to standard dynamic range (SDR) data in different situations. Moreover, we analyze the influence of texture information and pixel distribution of input data on the performance of the DNNs-based detector. Code and dataset will be available at https://gitee.com//mindspore/models/tree/master/research/cv/RAOD. Ruikang Xu, Chang Chen 0004, Jingyang Peng, Cheng Li 0009, Yibin Huang, Fenglong Song, Youliang Yan, Zhiwei Xiong |
CVPR | 7 |
| 2023 | PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video RetrievalabstractText-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video retrieval. However, due to the modality difference between videos and images, how to effectively adapt CLIP to the video domain is still underexplored. In this paper, we investigate this problem from two aspects. First, we enhance the transferred image encoder of CLIP for fine-grained video understanding in a seamless fashion. Second, we conduct fine-grained contrast between videos and texts from both model improvement and loss design. Particularly, we propose a fine-grained contrastive model equipped with parallel isomeric attention and dynamic routing, namely PIDRo, for text-video retrieval. The parallel isomeric attention module is used as the video encoder, which consists of two parallel branches modeling the spatial-temporal information of videos from both patch and frame levels. The dynamic routing module is constructed to enhance the text encoder of CLIP, generating informative word representations by distributing the fine-grained information to the related word tokens within a sentence. Such model design provides us with informative patch, frame and word representations. We then conduct token-wise interaction upon them. With the enhanced encoders and the token-wise loss, we are able to achieve finer-grained text-video alignment and more accurate retrieval. PIDRo obtains state-of-the-art performance over various text-video retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, DiDeMo and ActivityNet. Peiyan Guan, Renjing Pei, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu 0004, Songcen Xu, Youliang Yan, Edmund Y. Lam |
ICCV | 9 |
| 2023 | HiVLP: Hierarchical Interactive Video-Language Pre-TrainingabstractVideo-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and video-text pairs, which significantly outperforms pre-training with only the video-text pairs. Besides, existing methods usually model the cross-modal interaction using cross-attention between single-scale visual tokens and textual tokens. These visual features are either of low resolutions lacking fine-grained information, or of high resolutions without high-level semantics. To address the issue, we propose Hierarchical interactive Video-Language Pre-training (HiVLP) that efficiently uses a hierarchical visual feature group for multi-modal cross-attention during pre-training. In the hierarchical framework, low-resolution features are learned with focus on more global high-level semantic information, while high-resolution features carry fine-grained details. As a result, HiVLP has the ability to effectively learn both the global and fine-grained representations to achieve better alignment between video and text inputs. Furthermore, we design a hierarchical multi-scale vision contrastive loss for self-supervised learning to boost the interaction between them. Experimental results show that HiVLP establishes new state-of-the-art results in three downstream tasks, text-video retrieval, video-text retrieval, and video captioning. Jianzhuang Liu, Renjing Pei, Songcen Xu, Peng Dai 0002, Juwei Lu, Weimian Li, Youliang Yan |
ICCV | 8 |
| 2023 | Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the WildabstractThis paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependent texture, and a diffuse UV texture. Subsequently, users can either manually edit the UV or provide a prompt for the automatic generation of a new 3D-consistent texture. To achieve high-quality 3D texture generation, we propose a structure-aware score distillation sampling method to optimize a neural UV texture based on user-defined text and empower an image diffusion model with 3D-consistent generation capability. Furthermore, we introduce a few-view resampling training method and utilize a super-resolution model to obtain refined high-resolution UV textures (2048$\times$2048) for 3D texturing. Extensive experiments collectively validate the superior performance of Decorate3D in retexturing real-world 3D objects. Project page: https://decorate3d.github.io/Decorate3D/. Xinxin Zuo, Peng Dai 0002, Juwei Lu, Li Cheng 0001, Youliang Yan, Songcen Xu |
NeurIPS | 7 |
| 2022 | SiamTrans: Zero-Shot Multi-Frame Image Restoration with Pre-trained Siamese TransformersabstractWe propose a novel zero-shot multi-frame image restoration method for removing unwanted obstruction elements (such as rains, snow, and moire patterns) that vary in successive frames. It has three stages: transformer pre-training, zero-shot restoration, and hard patch refinement. Using the pre-trained transformers, our model is able to tell the motion difference between the true image information and the obstructing elements. For zero-shot image restoration, we design a novel model, termed SiamTrans, which is constructed by Siamese transformers, encoders, and decoders. Each transformer has a temporal attention layer and several self-attention layers, to capture both temporal and spatial information of multiple frames. Only self-supervisedly pre-trained on the denoising task, SiamTrans is tested on three different low-level vision tasks (deraining, demoireing, and desnowing). Compared with related methods, SiamTrans achieves the best performances, even outperforming those with supervised learning. Lin Liu 0016, Shanxin Yuan, Jianzhuang Liu, Youliang Yan, Qi Tian 0001 |
AAAI | 5 |
| 2022 | CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation
Zhihao Li 0002, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, Youliang Yan |
ECCV (5) | 5 |
| 2022 | A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution
Hengsheng Zhang, Xueyi Zou, Jiaming Guo, Youliang Yan, Rong Xie 0004, Li Song 0001 |
ECCV (17) | 4 |
| 2022 | Flow-Guided Sparse Transformer for Video DeblurringabstractExploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Transformer (FGST), for video deblurring. In FGST, we customize a self-attention module, Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA). For each $query$ element on the blurry reference frame, FGSW-MSA enjoys the guidance of the estimated optical flow to globally sample spatially sparse yet highly related $key$ elements corresponding to the same scene patch in neighboring frames. Besides, we present a Recurrent Embedding (RE) mechanism to transfer information from past frames and strengthen long-range temporal dependencies. Comprehensive experiments demonstrate that our proposed FGST outperforms state-of-the-art (SOTA) methods on both DVD and GOPRO datasets and yields visually pleasant results in real video deblurring. https://github.com/linjing7/VR-Baseline Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Youliang Yan, Xueyi Zou, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
ICML | 5 |
| 2022 | Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video RestorationabstractHow to properly model the inter-frame relation within the video sequence is an important but unsolved challenge for video restoration (VR). In this work, we propose an unsupervised flow-aligned sequence-to-sequence model (S2SVR) to address this problem. On the one hand, the sequence-to-sequence model, which has proven capable of sequence modeling in the field of natural language processing, is explored for the first time in VR. Optimized serialization modeling shows potential in capturing long-range dependencies among frames. On the other hand, we equip the sequence-to-sequence model with an unsupervised optical flow estimator to maximize its potential. The flow estimator is trained with our proposed unsupervised distillation loss, which can alleviate the data discrepancy and inaccurate degraded optical flow issues of previous flow-based methods. With reliable optical flow, we can establish accurate correspondence among multiple frames, narrowing the domain difference between 1D language and 2D misaligned frames and improving the potential of the sequence-to-sequence model. S2SVR shows superior performance in multiple VR tasks, including video deblurring, video super-resolution, and compressed video quality enhancement. https://github.com/linjing7/VR-Baseline Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Youliang Yan, Xueyi Zou, Yulun Zhang 0001, Luc Van Gool |
ICML | 5 |
| 2022 | Multi-Scale Coarse-to-Fine Transformer for Frame InterpolationabstractThe majority of prevailing video interpolation methods compute flows to estimate the intermediate motion. However, accurate estimation of the intermediate motion is difficult with low-order motion model hypothesis, which induces enormous difficulties for subsequent processing. To alleviate the limitation, we propose a two-stage flow-free video interpolation architecture. Rather than utilizing pre-defined motion models, our method represents complex motion through data-driven learning. In the first stage, we analyze spatial-temporal information and generate coarse anchor frame features. In the second stage, we employ transformers to transfer neighboring features to the intermediate time steps and enhance the spatial textures. To improve the quality of coarse anchor frame features and the robustness in dealing with the multi-scale textures with large-scale motion, we propose a multi-scale architecture and transformers with variable token sizes to progressively enhance the features. The experimental results demonstrate that our model outperforms state-of-the-art methods for both single frame and multi frames interpolation tasks, and the extended ablation studies verify the effectiveness of our model. Chen Li 0021, Li Song 0001, Xueyi Zou, Jiaming Guo, Youliang Yan, Wenjun Zhang 0001 |
ACM Multimedia | 5 |
| 2021 | Real-time Image Enhancer via Learnable Spatial-aware 3D Lookup TablesabstractRecently, deep learning-based image enhancement algorithms achieved state-of-the-art (SOTA) performance on several publicly available datasets. However, most existing methods fail to meet practical requirements either for visual perception or for computation efficiency, especially for high-resolution images. In this paper, we propose a novel real-time image enhancer via learnable spatial-aware 3dimentional lookup tables(3D LUTs), which well considers global scenario and local spatial information. Specifically, we introduce a light weight two-head weight predictor that has two outputs. One is a 1D weight vector used for image-level scenario adaptation, the other is a 3D weight map aimed for pixel-wise category fusion. We learn the spatial-aware 3D LUTs and fuse them according to the aforementioned weights in an end-to-end manner. The fused LUT is then used to transform the source image into the target tone in an efficient way. Extensive results show that our model outperforms SOTA image enhancement methods on public datasets both subjectively and objectively, and that our model only takes about 4ms to process a 4K resolution image on one NVIDIA V100 GPU. Yong Li 0008, Jingyang Peng, Yipeng Ma, Fenglong Song, Youliang Yan |
ICCV | 7 |
| 2020 | BlendMask: Top-Down Meets Bottom-Up for Instance SegmentationabstractInstance segmentation is one of the fundamental vision tasks. Recently, fully convolutional instance segmentation methods have drawn much attention as they are often simpler and more efficient than two-stage approaches like Mask R-CNN. To date, almost all such approaches fall behind the two-stage Mask R-CNN method in mask precision when models have similar computation complexity, leaving great room for improvement. In this work, we achieve improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity. Our main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches. The proposed BlendMask can effectively predict dense per-pixel position-sensitive instance features with very few channels, and learn attention maps for each instance with merely one convolution layer, thus being fast in inference. BlendMask can be easily incorporate with the state-of-the-art one-stage detection frameworks and outperforms Mask R-CNN under the same training schedule while being faster. A light-weight version of BlendMask achieves 36.0 mAP at 27 FPS evaluated on a single 1080Ti. Because of its simplicity and efficacy, we hope that our BlendMask could serve as a simple yet strong baseline for a wide range of instance-wise prediction tasks. Hao Chen 0041, Kunyang Sun, Zhi Tian, Chunhua Shen, Yongming Huang 0001, Youliang Yan |
CVPR | 6 |
| 2020 | Mask Encoding for Single Shot Instance SegmentationabstractTo date, instance segmentation is dominated by two-stage methods, as pioneered by Mask R-CNN. In contrast, one-stage alternatives cannot compete with Mask R-CNN in mask AP, mainly due to the difficulty of compactly representing masks, making the design of one-stage methods very challenging. In this work, we propose a simple single-shot instance segmentation framework, termed mask encoding based instance segmentation (MEInst). Instead of predicting the two-dimensional mask directly, MEInst distills it into a compact and fixed-dimensional representation vector, which allows the instance segmentation task to be incorporated into one-stage bounding-box detectors and results in a simple yet efficient instance segmentation framework. The proposed one-stage MEInst achieves 36.4% in mask AP with single-model (ResNeXt-101-FPN backbone) and single-scale testing on the MS-COCO benchmark. We show that the much simpler and flexible one-stage instance segmentation method, can also achieve competitive performance. This framework can be easily adapted for other instance-level recognition tasks. Code is available at: git.io/AdelaiDet Rufeng Zhang, Zhi Tian, Chunhua Shen, Mingyu You, Youliang Yan |
CVPR | 5 |
| 2019 | Knowledge Adaptation for Efficient Semantic SegmentationabstractBoth accuracy and efficiency are of significant importance to the task of semantic segmentation. Existing deep FCNs suffer from heavy computations due to a series of high-resolution feature maps for preserving the detailed knowledge in dense estimation. Although reducing the feature map resolution (i.e., applying a large overall stride) via subsampling operations (e.g., polling and convolution striding) can instantly increase the efficiency, it dramatically decreases the estimation accuracy. To tackle this dilemma, we propose a knowledge distillation method tailored for semantic segmentation to improve the performance of the compact FCNs with large overall stride. To handle the inconsistency between the features of the student and teacher network, we optimize the feature similarity in a transferred latent domain formulated by utilizing a pre-trained autoencoder. Moreover, an affinity distillation module is proposed to capture the long-range dependency by calculating the non local interactions across the whole image. To validate the effectiveness of our proposed method, extensive experiments have been conducted on three popular benchmarks: Pascal VOC, Cityscapes and Pascal Context. Built upon a highly competitive baseline, our proposed method can improve the performance of a student network by 2.5% (mIOU boosts from 70.2 to 72.7 on the cityscapes test set) and can train a better compact model with only 8% float operations (FLOPS) of a model that achieves comparable performances. Tong He 0001, Chunhua Shen, Zhi Tian, Dong Gong, Changming Sun, Youliang Yan |
CVPR | 6 |
| 2019 | Decoders Matter for Semantic Segmentation: Data-Dependent Decoding Enables Flexible Feature AggregationabstractRecent semantic segmentation methods exploit encoder-decoder architectures to produce the desired pixel-wise segmentation prediction. The last layer of the decoders is typically a bilinear upsampling procedure to recover the final pixel-wise prediction. We empirically show that this oversimple and data-independent bilinear upsampling may lead to sub-optimal results. In this work, we propose a data-dependent upsampling (DUpsampling) to replace bilinear, which takes advantages of the redundancy in the label space of semantic segmentation and is able to recover the pixel-wise prediction from low-resolution outputs of CNNs. The main advantage of the new upsampling layer lies in that with a relatively lower-resolution feature map such as 1/16 or 1/32 of the input size, we can achieve even better segmentation accuracy, significantly reducing computation complexity. This is made possible by 1) the new upsampling layer's much improved reconstruction capability; and more importantly 2) the DUpsampling based decoder's flexibility in leveraging almost arbitrary combinations of the CNN encoders' features. Experiments on PASCAL VOC demonstrate that with much less computation complexity, our decoder outperforms the state-of-the-art decoder. Finally, without any post-processing, the framework equipped with our proposed decoder achieves new state-of-the-art performance on two datasets: 88.1% mIOU on PASCAL VOC with 30% computation of the previously best model; and 52.5% mIOU on PASCAL Context. Zhi Tian, Tong He 0001, Chunhua Shen, Youliang Yan |
CVPR | 4 |
| 2019 | Enforcing Geometric Constraints of Virtual Normal for Depth PredictionabstractMonocular depth prediction plays a crucial role in understanding 3D scene geometry. Although recent methods have achieved impressive progress in evaluation metrics such as the pixel-wise relative error, most methods neglect the geometric constraints in the 3D space. In this work, we show the importance of the high-order 3D geometric constraints for depth prediction. By designing a loss term that enforces one simple type of geometric constraints, namely, virtual normal directions determined by randomly sampled three points in the reconstructed 3D space, we can considerably improve the depth prediction accuracy. Furthermore, we can not only predict accurate depth but also achieve high-quality other 3D information from the depth without retraining new parameters, Significantly, the byproduct of this predicted depth being sufficiently accurate is that we are now able to recover good 3D structures of the scene such as the point cloud and surface normal directly from the depth, eliminating the necessity of training new sub-models as was previously done. Experiments on two challenging benchmarks: NYU Depth-V2 and KITTI demonstrate the effectiveness of our method and state-of-the-art performance. Wei Yin 0006, Yifan Liu 0001, Chunhua Shen, Youliang Yan |
ICCV | 4 |
| 2019 | Exploiting Temporal Consistency for Real-Time Video Depth EstimationabstractAccuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation performance. In this work, we focus on exploring temporal information from monocular videos for depth estimation. Specifically, we take the advantage of convolutional long short-term memory (CLSTM) and propose a novel spatial-temporal CSLTM (ST-CLSTM) structure. Our ST-CLSTM structure can capture not only the spatial features but also the temporal correlations/consistency among consecutive video frames with negligible increase in computational cost. Additionally, in order to maintain the temporal consistency among the estimated depth frames, we apply the generative adversarial learning scheme and design a temporal consistency loss. The temporal consistency loss is combined with the spatial loss to update the model in an end-to-end fashion. By taking advantage of the temporal information, we build a video depth estimation framework that runs in real-time and generates visually pleasant results. Moreover, our approach is flexible and can be generalized to most existing depth estimation frameworks. Code is available at: https://tinyurl.com/STCLSTM Haokui Zhang, Ying Li 0017, Yuanzhouhan Cao, Yu Liu 0029, Chunhua Shen, Youliang Yan |
ICCV | 6 |
| 2016 | HiGene: A high-performance platform for genomic data analysisabstractPost-sequencing genomic data analysis becomes a major challenge while next-generation sequencing technologies evolve by leaps and bounds. The data-intensive and compute-intensive nature of genome analysis makes cluster computing an attractive choice for building efficient solutions. This paper presents HiGene, a high-performance genome analysis platform that exploits big data technology to revolutionize genomics data crunching power. HiGene reconstructs the genome analysis pipeline by exploiting both multi-core and multi-node parallelization using Apache Spark, and employs two key techniques to further boost the performance. First, a dynamic computing resource re-allocator is implemented, which allows flexible on-demand resource allocation for operations inside tasks. Second, an efficient skew mitigation approach is proposed, which automatically identifies and resolves data skew and computation skew through task repartitioning and resource reallocating respectively. HiGene has been evaluated with a whole human genome dataset on a 10-node Huawei 5885 cluster. Experimental results show that HiGene achieves remarkable high performance that reduces the total running time on a whole genome sequence dataset from days to nearly one hour. Furthermore, it is two times faster than state-of-the-art cluster based approaches. Liqun Deng, Guowei Huang 0002, Yuzheng Zhuang, Jiansheng Wei, Youliang Yan |
BIBM | 5 |
| 2015 | Interactive visual summary of major communities in a large networkabstractIn this paper, we introduce a novel visualization method which allows people to explore, compare and refine the major communities in a large network. We first detect major communities in a network using data mining and community analysis methods. Then, the statistics attributes of each community, the relational strength between communities, and the boundary nodes connecting those communities are computed and stored. We propose a novel method based on Voronoi treemap to encode each community with a polygon and the relative positions of polygons encode their relational strengths. Different community attributes can be encoded by polygon shapes, sizes and colors. A corner-cutting method is further introduced to adjust the smoothness of polygons based on certain community attribute. To accommodate the boundary nodes, the gaps between the polygons are widened by a polygon-shrinking algorithm such that the boundary nodes can be conveniently embedded into the newly created spaces. The method is very efficient, enabling users to test different community detection algorithms, fine tune the results, and explore the fuzzy relations between communities interactively. The case studies with two real data sets demonstrate that our approach can provide a visual summary of major communities in a large network, and help people better understand the characteristics of each community and inspect various relational patterns between communities. Sixiao Yang, Youliang Yan, Huamin Qu |
PacificVis | 4 |
| 2015 | Hadoop+: Modeling and Evaluating the Heterogeneity for MapReduce Applications in Heterogeneous ClustersabstractDespite the widespread adoption of heterogeneous clusters in modern data centers, modeling heterogeneity is still a big challenge, especially for large-scale MapReduce applications. In a CPU/GPU hybrid heterogeneous cluster, allocating more computing resources to a MapReduce application does not always mean better performance, since simultaneously running CPU and GPU tasks will contend for shared resources. Wenting He, Huimin Cui, Binbin Lu, Shengmei Li, Gong Ruan, Jingling Xue, Xiaobing Feng 0002, Wensen Yang, Youliang Yan |
ICS | 10 |
| 2014 | Rhymes: A shared virtual memory system for non-coherent tiled many-core architecturesabstractThe rising core count per processor is pushing chip complexity to a level that hardware-based cache coherency protocols become too hard and costly to scale someday. We need new designs of many-core hardware and software other than traditional technologies to keep up with the ever-increasing scalability demands. A cluster-on-chip architecture, as exemplified by the Intel Single-chip Cloud Computer (SCC), promotes a software-oriented approach instead of hardware support to implementing shared memory coherence. This paper presents a shared virtual memory (SVM) system, dubbed Rhymes, tailored to new processor kinds of non-coherent and hybrid memory architectures. Rhymes features a two-way cache coherence protocol to enforce release consistency for pages allocated in shared physical memory (SPM) and scope consistency for pages in percore private memory. It also supports page remapping on a percore basis to boost data locality. We implement and test Rhymes on the SCC port of the Barrelfish OS. Experimental results show that our SVM outperforms the pure SPM approach used by Intel's software managed coherence (SMC) library by up to 12 times through improved cache utilization for applications with strong data reuse patterns. King Tin Lam, Jinghao Shi, Dominic Hung, Cho-Li Wang, Zhiquan Lai, Wangbin Zhu, Youliang Yan |
ICPADS | 7 |
| 2013 | An empirical model for predicting cross-core performance interference on multicore processorsabstractDespite their widespread adoption in cloud computing, multicore processors are heavily under-utilized in terms of computing resources. To avoid the potential for negative and unpredictable interference, co-location of a latency-sensitive application with others on the same multicore processor is disallowed, leaving many cores idle and causing low machine utilization. To enable co-location while providing QoS guarantees, it is challenging but important to predict performance interference between co-located applications. This research is driven by two key insights. First, the performance degradation of an application can be represented as a predictor function of the aggregate pressures on shared resources from all cores, regardless of which applications are co-running and what their individual pressures are. Second, a predictor function is piecewise rather than non-piecewise as in prior work, thereby enabling different types of dominant contention factors to be more accurately captured by different subfunctions in its different subdomains. Based on these insights, we propose to adopt a two-phase regression approach to efficiently building a predictor function. Validation using a large number of benchmarks and nine real-world datacenter applications on three different platforms shows that our approach is also precise, with an average error not exceeding 0.4%. When applied to the nine datacenter applications, our approach improves overall resource utilization from 50% to 88% at the cost of 10% QoS degradation. Xiaobing Feng 0002, Huimin Cui, Youliang Yan, Jingling Xue, Wensen Yang |
PACT | 4 |