VLDB 2026 Research / reviewers in the wild / expert
Wenxiu Sun
dblp:16/9879
· DBLP profile ↗
71ranked-venue papers
7as first author
27since 2021 · last 2025
0000-0001-5026-8820ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 37 · 18 since 2021Systems, architecture and hardware · 6 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing HDR Imaging with Joint Denoising and DeblurringabstractAbstract Considerable progress has been made in high dynamic range (HDR) image reconstruction from multi-exposure low dynamic range (LDR) frames in recent years. Despite achieving satisfactory HDR results while handling the misalignment among LDR frames, previous studies pay less attention to a prevalent but more crucial concern: the significant noise and motion blur that exist in multi-exposure frames captured by a handheld camera. To overcome these extensive and challenging corruptions, we achieve the HDR imaging task from two key aspects: First, due to the absence of related datasets, previous learning-based methods struggle with performing high-quality HDR imaging when the input LDR frames are confronted with real-world image noise and motion blur. Recognizing the importance of this aspect, we propose the first available realistic dataset based on the real-world burst imaging pipeline for training and evaluating different methods in the joint HDR imaging, denoising, and deblurring task. Second, due to the corruption-insensitive of previous network architectures, we propose a novel and efficient attention-based multi-exposure HDR imaging method, which skillfully selects the optimal information (clean or sharp) from corrupted LDR inputs by our customized cross-attention mechanism to generate HDR information. Furthermore, to enhance the robustness of our cross-attention mechanism, we introduce a novel E ntropy D ecreasing loss (ED loss) which decreases the entropy of the calculated attention map to alleviate the ghosting artifacts during multi-exposure information fusion. Extensive experimental results demonstrate that the proposed method trained on the proposed dataset surpasses related state-of-the-art methods with outstanding real-world photography quality. Codes and the dataset will be available. Zhefan Rao, Chenyang Lei, Wenxiu Sun, Qiong Yan, Qifeng Chen 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Deep Unrolled Weighted Graph Laplacian Regularization for Depth Completion
Jin Zeng 0004, Qingpeng Zhu, Tongxuan Tian, Wenxiu Sun, Lin Zhang 0014, Shengjie Zhao 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Iterative Token Evaluation and Refinement for Real-World Super-resolutionabstractReal-world image super-resolution (RWSR) is a long-standing problem as low-quality (LQ) images often have complex and unidentified degradations. Existing methods such as Generative Adversarial Networks (GANs) or continuous diffusion models present their own issues including GANs being difficult to train while continuous diffusion models requiring numerous inference steps. In this paper, we propose an Iterative Token Evaluation and Refinement (ITER) framework for RWSR, which utilizes a discrete diffusion model operating in the discrete token representation space, i.e., indexes of features extracted from a VQGAN codebook pre-trained with high-quality (HQ) images. We show that ITER is easier to train than GANs and more efficient than continuous diffusion models. Specifically, we divide RWSR into two sub-tasks, i.e., distortion removal and texture generation. Distortion removal involves simple HQ token prediction with LQ images, while texture generation uses a discrete diffusion model to iteratively refine the distortion removal output with a token refinement network. In particular, we propose to include a token evaluation network in the discrete diffusion process. It learns to evaluate which tokens are good restorations and helps to improve the iterative refinement results. Moreover, the evaluation network can first check status of the distortion removal output and then adaptively select total refinement steps needed, thereby maintaining a good balance between distortion removal and texture generation. Extensive experimental results show that ITER is easy to train and performs well within just 8 iterative steps. Chaofeng Chen, Shangchen Zhou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin |
AAAI | 5 |
| 2024 | Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation ModelsabstractMulti-modality large language models (MLLMs), as represented by GPT-4V, have introduced a paradigm shift for visual perception and understanding tasks, that a variety of abilities can be achieved within one foundation model. While current MLLMs demonstrate primary low-level visual abilities from the identification of low-level visual attributes (e.g., clarity, brightness) to the evaluation on image quality, there's still an imperative to further improve the accuracy of MLLMs to substantially alleviate human burdens. To address this, we collect the first dataset consisting of human natural language feedback on low-level vision. Each feedback offers a comprehensive description of an image's low-level visual attributes, culminating in an overall quality assessment. The constructed Q-Pathway dataset includes 58K detailed human feedbacks on 18,973 multi-sourced images with diverse low-level appearance. To ensure MLLMs can adeptly handle diverse queries, we further propose a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct. Experimental results indicate that the Q-Instruct consistently elevates various low-level visual capabilities across multiple base models. We anticipate that our datasets can pave the way for a future that foundation models can assist humans on low-level visual tasks. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Kaixin Xu, Chunyi Li 0001, Jingwen Hou, Guangtao Zhai, Geng Xue, Wenxiu Sun, Qiong Yan, Weisi Lin |
CVPR | 12 |
| 2024 | Enhancing Diffusion Models with Text-Encoder Reinforcement Learning
Chaofeng Chen, Annan Wang, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin |
ECCV (25) | 5 |
| 2024 | Towards Open-Ended Visual Quality Comparison
Haoning Wu 0001, Hanwei Zhu, Erli Zhang 0001, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Xiaohong Liu 0001, Guangtao Zhai, Shiqi Wang 0001, Weisi Lin |
ECCV (3) | 9 |
| 2024 | Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionabstractThe rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in assessing the abilities of MLLMs on **low-level visual perception and understanding**. To address this gap, we present **Q-Bench**, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment. **_a)_** To evaluate the low-level **_perception_** ability, we construct the **LLVisionQA** dataset, consisting of 2,990 diverse-sourced images, each equipped with a human-asked question focusing on its low-level attributes. We then measure the correctness of MLLMs on answering these questions. **_b)_** To examine the **_description_** ability of MLLMs on low-level information, we propose the **LLDescribe** dataset consisting of long expert-labelled *golden* low-level text descriptions on 499 images, and a GPT-involved comparison pipeline between outputs of MLLMs and the *golden* descriptions. **_c)_** Besides these two tasks, we further measure their visual quality **_assessment_** ability to align with human opinion scores. Specifically, we design a softmax-based strategy that enables MLLMs to predict *quantifiable* quality scores, and evaluate them on various existing image quality assessment (IQA) datasets. Our evaluation across the three abilities confirms that MLLMs possess preliminary low-level visual skills. However, these skills are still unstable and relatively imprecise, indicating the need for specific enhancements on MLLMs towards these abilities. We hope that our benchmark can encourage the research community to delve deeper to discover and enhance these untapped potentials of MLLMs. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Annan Wang, Chunyi Li 0001, Wenxiu Sun, Qiong Yan, Guangtao Zhai, Weisi Lin |
ICLR | 8 |
| 2024 | Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsabstractThe explosion of visual content available online underscores the requirement for an accurate machine assessor to robustly evaluate scores across diverse types of visual contents. While recent studies have demonstrated the exceptional potentials of large multi-modality models (LMMs) on a wide range of related fields, in this work, we explore how to teach them for visual rating aligning with human opinions. Observing that human raters only learn and judge discrete text-defined levels in subjective studies, we propose to emulate this subjective process and teach LMMs with text-defined rating levels instead of scores. The proposed Q-Align achieves state-of-the-art accuracy on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) under the original LMM structure. With the syllabus, we further unify the three tasks into one model, termed the OneAlign. Our experiments demonstrate the advantage of discrete levels over direct scores on training, and that LMMs can learn beyond the discrete levels and provide effective finer-grained evaluations. Code and weights will be released. Haoning Wu 0001, Weixia Zhang, Chaofeng Chen, Chunyi Li 0001, Annan Wang, Erli Zhang 0001, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, Weisi Lin |
ICML | 10 |
| 2024 | Q-Ground: Image Quality Grounding with Large Multi-modality ModelsabstractRecent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual content. However, these advancements are mostly focused on overall quality assessment, and the detailed examination of local quality, which is crucial for comprehensive visual understanding, is still largely unexplored. In this work, we introduce Q-Ground, the first framework aimed at tackling fine-scale visual quality grounding by combining large multi-modality models with detailed visual quality analysis. Cen- tral to our contribution is the introduction of the QGround-100K dataset, a novel resource containing 100k triplets of (image, quality text, distortion segmentation) to facilitate deep investigations into visual quality. The dataset comprises two parts: one with human- labeled annotations for accurate quality assessment, and another la- beled automatically by LMMs such as GPT4V, which helps improve the robustness of model training while also reducing the costs of data collection. With the QGround-100K dataset, we propose a LMM-based method equipped with multi-scale feature learning to learn models capable of performing both image quality answer- ing and distortion segmentation based on text prompts. This dual- capability approach not only refines the model’s understanding of region-aware image quality but also enables it to interactively re- spond to complex, text-based queries about image quality and spe- cific distortions. Q-Ground takes a step towards sophisticated vi- sual quality analysis in a finer scale, establishing a new benchmark for future research in the area. Codes and dataset are available at https://github.com/Q-Future/Q-Ground. Chaofeng Chen, Sensen Yang, Haoning Wu 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ACM Multimedia | 7 |
| 2024 | TOPIQ: A Top-Down Approach From Semantics to Distortions for Image Quality AssessmentabstractImage Quality Assessment (IQA) is a fundamental task in computer vision that has witnessed remarkable progress with deep neural networks. Inspired by the characteristics of the human visual system, existing methods typically use a combination of global and local representations (i.e., multi-scale features) to achieve superior performance. However, most of them adopt simple linear fusion of multi-scale features, and neglect their possibly complex relationship and interaction. In contrast, humans typically first form a global impression to locate important regions and then focus on local details in those regions. We therefore propose a top-down approach that uses high-level semantics to guide the IQA network to focus on semantically important local distortion regions, named as TOPIQ. Our approach to IQA involves the design of a heuristic coarse-to-fine network (CFANet) that leverages multi-scale features and progressively propagates multi-level semantic information to low-level representations in a top-down manner. A key component of our approach is the proposed cross-scale attention mechanism, which calculates attention maps for lower level features guided by higher level features. This mechanism emphasizes active semantic regions for low-level distortions, thereby improving performance. TOPIQ can be used for both Full-Reference (FR) and No-Reference (NR) IQA. We use ResNet50 as its backbone and demonstrate that TOPIQ achieves better or competitive performance on most public FR and NR benchmarks compared with state-of-the-art methods based on vision transformers, while being much more efficient (with only ∼ 13% FLOPS of the current best FR method). Codes are released at https://github.com/chaofengc/IQA-PyTorch. Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu 0001, Wenxiu Sun, Qiong Yan, Weisi Lin |
IEEE Trans. Image Process. | 6 |
| 2024 | Blind Video Quality Prediction by Uncovering Human Video Perceptual RepresentationabstractBlind video quality assessment (VQA) has become an increasingly demanding problem in automatically assessing the quality of ever-growing in-the-wild videos. Although efforts have been made to measure temporal distortions, the core to distinguish between VQA and image quality assessment (IQA), the lack of modeling of how the human visual system (HVS) relates to the temporal quality of videos hinders the precise mapping of predicted temporal scores to the human perception. Inspired by the recent discovery of the temporal straightness law of natural videos in the HVS, this paper intends to model the complex temporal distortions of in-the-wild videos in a simple and uniform representation by describing the geometric properties of videos in the visual perceptual domain. A novel videolet, with perceptual representation embedding of a few consecutive frames, is designed as the basic quality measurement unit to quantify temporal distortions by measuring the angular and linear displacements from the straightness law. By combining the predicted score on each videolet, a perceptually temporal quality evaluator (PTQE) is formed to measure the temporal quality of the entire video. Experimental results demonstrate that the perceptual representation in the HVS is an efficient way of predicting subjective temporal quality. Moreover, when combined with spatial quality metrics, PTQE achieves top performance over popular in-the-wild video datasets. More importantly, PTQE requires no additional information beyond the video being assessed, making it applicable to any dataset without parameter tuning. Additionally, the generalizability of PTQE is evaluated on video frame interpolation tasks, demonstrating its potential to benefit temporal-related enhancement tasks. Kangmin Xu, Haoning Wu 0001, Chaofeng Chen, Wenxiu Sun, Qiong Yan, C.-C. Jay Kuo, Weisi Lin |
IEEE Trans. Image Process. | 5 |
| 2023 | Memory-Efficient and Real-Time SPAD-based dToF Depth Sensor with Spatial and Statistical CorrelationabstractSingle Photon Avalanche Diode (SPAD)-based direct time-of-flight (dToF) depth sensors are widely used in Internet of Things (IoT) devices due to their high accuracy. Existing SPAD-based dToF sensors measure depth by continually accumulating the depth-measured value in a histogram. However, histogram-based methods typically have low convergence speed (~10 frames per second (FPS)) and large memory overhead (MB-level), hindering their use in real-time embedded IoT devices. To overcome these two challenges, we propose SSC, a histogram-free Spatial and Statistical Correlation based depth measurement method. On the one hand, SSC applies the spatial correlation of the adjacent pixels to accelerate the convergence speed. On the other hand, SSC explores the statistical correlation of depth measurements to reduce the memory overhead. In order to implement SSC with small hardware area and low power, we design mert-dToF, a memory-efficient and real-time dToF sensor for efficient execution. mert-dToF abstracts mainly operations in SSC into four basic operators and designs corresponding hardware with a fine-grained pipeline to maximize resource reuse and computational parallelism. Extensive experiments show that compared with state-of-the-art (SOTA) histogram-based dToF sensors, mert-dToF achieves ~8% accuracy improvement and 7.80× speedup (from 6.24 FPS to 48.70 FPS). The memory overhead is reduced by up to 60.91% (from 48 KB to 18.75 KB). Zhenhua Zhu 0002, Qingpeng Zhu, Jiangwei Zhang, Wenxiu Sun, Guohao Dai 0001, Fei Qiao, Huazhong Yang, Yu Wang 0002 |
DAC | 6 |
| 2023 | Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesabstractThe rapid increase in user-generated content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the $\color{Green}{\text{technical perspective}}$, measuring the perception of distortions; and the $\color{Blue}{\text{aesthetic perspective}}$, which relates to preference and recommendation on contents. To understand how these two perspectives affect overall subjective opinions in UGC-VQA, we conduct a large-scale subjective study to collect human quality opinions on the overall quality of videos as well as perceptions from aesthetic and technical perspectives. The collected Disentangled Video Quality Database (DIVIDE-3k) confirms that human quality opinions on UGC videos are universally and inevitably affected by both aesthetic and technical perspectives. In light of this, we propose the Disentangled Objective Video Quality Evaluator (DOVER) to learn the quality of UGC videos based on the two perspectives. The DOVER proves state-of-the-art performance in UGC-VQA under very high efficiency. With perspective opinions in DIVIDE-3k, we further propose DOVER++, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective. Code at https://github.com/VQAssessment/DOVER. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ICCV | 7 |
| 2023 | Exploring Opinion-Unaware Video Quality Assessment with Semantic Affinity CriterionabstractRecent learning-based video quality assessment (VQA) algorithms are expensive to implement due to the cost of data collection of human quality opinions, and are less robust across various scenarios due to the biases of these opinions. This motivates our exploration on opinion-unaware (a.k.a zero-shot) VQA approaches. Existing approaches only considers low-level naturalness in spatial or temporal domain, without considering impacts from high-level semantics. In this work, we introduce an explicit semantic affinity index for opinion-unaware VQA using text-prompts in the contrastive language-image pre-training (CLIP) model. We also aggregate it with different traditional low-level naturalness indexes through gaussian normalization and sigmoid rescaling strategies. Composed of aggregated semantic and technical metrics, the proposed Blind Unified Opinion-Unaware Video Quality Index via Semantic and Technical Metric Aggregation (BUONA-VISTA) outperforms existing opinion-unaware VQA methods by at least 20% improvements, and is more robust than opinion-aware approaches. Haoning Wu 0001, Jingwen Hou, Chaofeng Chen, Erli Zhang 0001, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ICME | 7 |
| 2023 | A Three-Step Multi-Resolution Time-to-Digital ConverterabstractThis work proposes a three-step multi-resolution time-to-digital converter (TDC) architecture based on the vernier delay line (VDL). The proposed architecture uses a delay-locked loop (DLL) to control TDC with a smooth coarse-to-fine strategy. In addition, the fine TDC uses a combination of multiple resolutions to reduce the number of delay cells and flip-flops. This architecture helps to reduce the area and power consumption and maintains high resolution. We proposed architecture performs better trade-offs between power consumption, linearity, accuracy, and measurement range. The simulation results show that the 7-bit TDC based on VDL designed in 180 nm CMOS achieves 5 ps of time resolution, 0.76/-0.8 LSB DNL and 1.02/-1.39 LSB INL at 100 MHz clock frequency while consuming 3.1 mW, which corresponds to the figure of merit (FoM) of 0.242 pJ/Conv. Jiang Yan, Yu Wang 0002, Fei Qiao, Jiangwei Zhang, Qi Wei 0001, Qingpeng Zhu, Wenxiu Sun, Ge Shi 0001 |
ISCAS | 9 |
| 2023 | Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted ApproachabstractThe proliferation of in-the-wild videos has greatly expanded the Video Quality Assessment (VQA) problem. Unlike early definitions that usually focus on limited distortion types, VQA on in-the-wild videos is especially challenging as it could be affected by complicated factors, including various distortions and diverse contents. Though subjective studies have collected overall quality scores for these videos, how the abstract quality scores relate with specific factors is still obscure, hindering VQA methods from more concrete quality evaluations (e.g. sharpness of a video). To solve this problem, we collect over two million opinions on 4,543 in-the-wild videos on 13 dimensions of quality-related factors, including in-capture authentic distortions (e.g. motion blur, noise, flicker), errors introduced by compression and transmission, and higher-level experiences on semantic contents and aesthetic issues (e.g. composition, camera trajectory), to establish the multi-dimensional Maxwell database. Specifically, we ask the subjects to label among a positive, a negative, and a neutral choice for each dimension. These explanation-level opinions allow us to measure the relationships between specific quality factors and abstract subjective quality ratings, and to benchmark different categories of VQA algorithms on each dimension, so as to more comprehensively analyze their strengths and weaknesses. Furthermore, we propose the MaxVQA, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in our analyses. The MaxVQA can jointly evaluate various specific quality factors and final quality scores with state-of-the-art accuracy on all dimensions, and superb generalization ability on existing datasets. Code and data available at https://github.com/VQAssessment/MaxVQA. Haoning Wu 0001, Erli Zhang 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ACM Multimedia | 7 |
| 2023 | Neighbourhood Representative Sampling for Efficient End-to-End Video Quality AssessmentabstractThe increased resolution of real-world videos presents a dilemma between efficiency and accuracy for deep Video Quality Assessment (VQA). On the one hand, keeping the original resolution will lead to unacceptable computational costs. On the other hand, existing practices, such as resizing or cropping, will change the quality of original videos due to difference in details or loss of contents, and are henceforth harmful to quality assessment. With obtained insight from the studies of spatial-temporal redundancy in the human visual system, visual quality around a neighbourhood has high probability to be similar, and this motivates us to investigate an effective quality-sensitive neighbourhood representative sampling scheme for VQA. In this work, we propose a unified scheme, spatial-temporal grid mini-cube sampling (St-GMS), and the resultant samples are namedfragments. In St-GMS, full-resolution videos are first divided into mini-cubes with predefined spatial-temporal grids, then the temporal-aligned quality representatives are sampled to compose the fragments that serve as inputs for VQA. In addition, we design the Fragment Attention Network (FANet), a network architecture tailored specifically for fragments. With fragments and FANet, the proposedFAST-VQAandFasterVQA(with an improved sampling scheme) achieves up to 1612× efficiency than the existing state-of-the-art, meanwhile achieving significantly better performance on all relevant VQA benchmarks. Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Jinwei Gu, Weisi Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | DisCoVQA: Temporal Distortion-Content Transformers for Video Quality AssessmentabstractCompared with spatial counterparts, temporal relationships between frames and their influences on video quality assessment (VQA) are still relatively under-studied in existing works. These relationships lead to two important types of effects for video quality. Firstly, some meaningless temporal variations (such as shaking, flicker, and unsmooth scene transitions) cause temporal distortions that degrade quality of videos. Secondly, the human visual system often has different attention to frames with different contents, resulting in their different importance to the overall video quality. Based on prominent time-series modeling ability of transformers, we propose a novel and effective transformer-based VQA method to tackle these two issues. To better differentiate temporal variations and thus capture the temporal distortions, we design the Spatial-Temporal Distortion Extraction (STDE) module that extracts multi-level spatial-temporal features with a video swin transformer tiny (Swin-T) backbone and uses temporal difference layer to further capture these distortions. To tackle with temporal quality attention, we propose the encoder-decoder-like temporal content transformer (TCT). We also introduce the temporal sampling on features to reduce the input length for the TCT, so as to improve the learning effectiveness and efficiency of this module. Consisting of the STDE and the TCT, the proposed Temporal Distortion-Content Transformers for Video Quality Assessment (DisCoVQA) reaches state-of-the-art performance on several VQA benchmarks without any extra pre-training datasets and up to 10% better generalization ability than existing methods. We also conduct extensive ablation experiments to prove the effectiveness of each part in our proposed model, and provide visualizations to prove that the proposed modules achieve our intention on modeling these temporal issues. Our code is published athttps://github.com/QualityAssessment/DisCoVQA. Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Wenxiu Sun, Qiong Yan, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling
Haoning Wu 0001, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, Weisi Lin |
ECCV (6) | 6 |
| 2022 | Exploring the Effectiveness of Video Perceptual Representation in Blind Video Quality AssessmentabstractWith the rapid growth of in-the-wild videos taken by non-specialists, blind video quality assessment (VQA) has become a challenging and demanding problem. Although lots of efforts have been made to solve this problem, it remains unclear how the human visual system (HVS) relates to the temporal quality of videos. Meanwhile, recent work has found that the frames of natural video transformed into the perceptual domain of the HVS tend to form a straight trajectory of the representations. With the obtained insight that distortion impairs the perceived video quality and results in a curved trajectory of the perceptual representation, we propose a temporal perceptual quality index (TPQI) to measure the temporal distortion by describing the graphic morphology of the representation. Specifically, we first extract the video perceptual representations from the lateral geniculate nucleus (LGN) and primary visual area (V1) of the HVS, and then measure the straightness and compactness of their trajectories to quantify the degradation in naturalness and content continuity of video. Experiments show that the perceptual representation in the HVS is an effective way of predicting subjective temporal quality, and thus TPQI can, for the first time, achieve comparable performance to the spatial quality metric and be even more effective in assessing videos with large temporal variations. We further demonstrate that by combining with NIQE, a spatial quality metric, TPQI can achieve top performance over popular in-the-wild video datasets. More importantly, TPQI does not require any additional information beyond the video being evaluated and thus can be applied to any datasets without parameter tuning. Source code is available at https://github.com/UoLMM/TPQI-VQA. Kangmin Xu, Haoning Wu 0001, Chaofeng Chen, Wenxiu Sun, Qiong Yan, Weisi Lin |
ACM Multimedia | 5 |
| 2022 | Exploiting Raw Images for Real-Scene Super-ResolutionabstractSuper-resolution is a fundamental problem in computer vision which aims to overcome the spatial limitation of camera sensors. While significant progress has been made in single image super-resolution, most algorithms only perform well on synthetic data, which limits their applications in real scenarios. In this paper, we study the problem of real-scene single image super-resolution to bridge the gap between synthetic data and real captured images. We focus on two issues of existing super-resolution algorithms: lack of realistic training data and insufficient utilization of visual information obtained from cameras. To address the first issue, we propose a method to generate more realistic training data by mimicking the imaging process of digital cameras. For the second issue, we develop a two-branch convolutional neural network to exploit the radiance information originally-recorded in raw images. In addition, we propose a dense channel-attention block for better image restoration as well as a learning-based guided filter network for effective color correction. Our model is able to generalize to different cameras without deliberately training on images from specific camera types. Extensive experiments demonstrate that the proposed algorithm can recover fine details and clear structures, and achieve high-quality results for single image super-resolution in real scenes. Xiangyu Xu 0002, Yongrui Ma, Wenxiu Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Deep Animation Video Interpolation in the WildabstractIn the animation industry, cartoon videos are usually produced at low frame rate since hand drawing of such frames is costly and time-consuming. Therefore, it is desirable to develop computational models that can automatically interpolate the in-between animation frames. However, existing video interpolation methods fail to produce satisfying results on animation data. Compared to natural videos, animation videos possess two unique characteristics that make frame interpolation difficult: 1) cartoons comprise lines and smooth color pieces. The smooth areas lack textures and make it difficult to estimate accurate motions on animation videos. 2) cartoons express stories via exaggeration. Some of the motions are non-linear and extremely large. In this work, we formally define and study the animation video interpolation problem for the first time. To address the aforementioned challenges, we propose an effective framework, AnimeInterp, with two dedicated modules in a coarse-to-fine manner. Specifically, 1) Segment-Guided Matching resolves the "lack of textures" challenge by exploiting global matching among color pieces that are piece-wise coherent. 2) Recurrent Flow Refinement resolves the "non-linear and extremely large motion" challenge by recur-rent predictions using a transformer-like architecture. To facilitate comprehensive training and evaluations, we build a large-scale animation triplet dataset, ATD-12K, which comprises 12,000 triplets with rich annotations. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art interpolation methods for animation videos. Notably, AnimeInterp shows favorable perceptual quality and robustness for animation scenarios in the wild. The proposed dataset and code are available at https://github.com/lisiyao21/AnimeInterp/. Li Siyao, Shiyu Zhao 0001, Weijiang Yu, Wenxiu Sun, Dimitris N. Metaxas, Chen Change Loy, Ziwei Liu 0002 |
CVPR | 4 |
| 2021 | Efficient Regional Memory Network for Video Object SegmentationabstractRecently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global matching between the current and past frames, which lead to mismatching to similar objects and high computational complexity. To address these problems, we propose a novel local-to-local matching solution for semi-supervised VOS, namely Regional Memory Network (RMNet). In RMNet, the precise regional memory is constructed by memorizing local regions where the target objects appear in the past frames. For the current query frame, the query regions are tracked and predicted based on the optical flow estimated from the previous frame. The proposed local-to-local matching effectively alleviates the ambiguity of similar objects in both memory and query frames, which allows the information to be passed from the regional memory to the query region efficiently and effectively. Experimental results indicate that the proposed RM-Net performs favorably against state-of-the-art methods on the DAVIS and YouTube-VOS datasets. Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Wenxiu Sun |
CVPR | 5 |
| 2021 | Dual-Camera Super-Resolution with Aligned Attention ModulesabstractWe present a novel approach to reference-based super-resolution (RefSR) with the focus on dual-camera super-resolution (DCSR), which utilizes reference images for high-quality and high-fidelity results. Our proposed method generalizes the standard patch-based feature matching with spatial alignment operations. We further explore the dual-camera super-resolution that is one promising application of RefSR, and build a dataset that consists of 146 image pairs from the main and telephoto cameras in a smartphone. To bridge the domain gaps between real-world images and the training images, we propose a self-supervised domain adaptation strategy for real-world images. Extensive experiments on our dataset and a public benchmark demonstrate clear improvement achieved by our method over state of the art in both quantitative evaluation and visual comparisons. Our code and data are available at https://tengfei-wang.github.io/Dual-Camera-SR/index.html. Tengfei Wang 0002, Jiaxin Xie, Wenxiu Sun, Qiong Yan, Qifeng Chen 0001 |
ICCV | 3 |
| 2021 | FuseFormer: Fusing Fine-Grained Information in Transformers for Video InpaintingabstractTransformer, as a strong and flexible architecture for modelling long-range relations, has been widely explored in vision tasks. However, when used in video inpainting that requires fine-grained representation, existed method still suffers from yielding blurry edges in detail due to the hard patch splitting. Here we aim to tackle this problem by proposing FuseFormer, a Transformer model designed for video inpainting via fine-grained feature fusion based on novel Soft Split and Soft Composition operations. The soft split divides feature map into many patches with given overlapping interval. On the contrary, the soft composition operates by stitching different patches into a whole feature map where pixels in overlapping regions are summed up. These two modules are first used in tokenization before Transformer layers and de-tokenization after Transformer layers, for effective mapping between tokens and features. Therefore, sub-patch level information interaction is enabled for more effective feature propagation between neighboring patches, resulting in synthesizing vivid content for hole regions in videos. Moreover, in FuseFormer, we elaborately insert the soft composition and soft split into the feed-forward network, enabling the 1D linear layers to have the capability of modelling 2D structure. And, the sub-patch level feature fusion ability is further enhanced. In both quantitative and qualitative evaluations, our proposed FuseFormer surpasses state-of-the-art methods. We also conduct detailed analysis to examine its superiority. Code and pretrained models are available at https://github.com/ruiliu-ai/FuseFormer. Rui Liu 0019, Hanming Deng, Yangyi Huang, Xiaoyu Shi 0002, Lewei Lu, Wenxiu Sun, Xiaogang Wang 0001, Jifeng Dai, Hongsheng Li 0001 |
ICCV | 6 |
| 2021 | Learning N: M Fine-grained Structured Sparse Neural Networks From Scratch
Aojun Zhou, Junnan Zhu, Wenxiu Sun, Hongsheng Li 0001 |
ICLR | 7 |
| 2021 | Toward 3D object reconstruction from stereo images
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Xiaojun Tong, Wenxiu Sun |
Neurocomputing | 6 |
| 2020 | Towards Geometry Guided Neural Relighting with Flash PhotographyabstractPrevious image based relighting methods require capturing multiple images to acquire high frequency lighting effect under different lighting conditions, which needs nontrivial effort and may be unrealistic in certain practical use scenarios. While such approaches rely entirely on cleverly sampling the color images under different lighting conditions, little has been done to utilize geometric information that crucially influences the high-frequency features in the images, such as glossy highlight and cast shadow. We therefore propose a framework for image relighting from a single flash photograph with its corresponding depth map using deep learning. By incorporating the depth map, our approach is able to extrapolate realistic high-frequency effects under novel lighting via geometry guided image decomposition from the flashlight image, and predict the cast shadow map from the shadow-encoding transformed depth map. Moreover, the single-image based setup greatly simplifies the data capture process. We experimentally validate the advantage of our geometry guided approach over state-of-the-art image-based approaches in intrinsic image decomposition and image relighting, and also demonstrate our performance on real mobile phone photo examples. Di Qiu, Jin Zeng 0004, Zhanghan Ke, Wenxiu Sun, Chengxi Yang |
3DV | 4 |
| 2020 | Polarized Reflection Removal With Perfect Alignment in the WildabstractWe present a novel formulation to removing reflection from polarized images in the wild. We first identify the misalignment issues of existing reflection removal datasets where the collected reflection-free images are not perfectly aligned with input mixed images due to glass refraction. Then we build a new dataset with more than 100 types of glass in which obtained transmission images are perfectly aligned with input mixed images. Second, capitalizing on the special relationship between reflection and polarized light, we propose a polarized reflection removal model with a two-stage architecture. In addition, we design a novel perceptual NCC loss that can improve the performance of reflection removal and general image decomposition tasks. We conduct extensive experiments, and results suggest that our model outperforms state-of-the-art methods on reflection removal. Chenyang Lei, Xuhua Huang, Qiong Yan, Wenxiu Sun, Qifeng Chen 0001 |
CVPR | 5 |
| 2020 | StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo MatchingabstractLarge-scale synthetic datasets are beneficial to stereo matching but usually introduce known domain bias. Although unsupervised image-to-image translation networks represented by CycleGAN show great potential in dealing with domain gap, it is non-trivial to generalize this method to stereo matching due to the problem of pixel distortion and stereo mismatch after translation. In this paper, we propose an end-to-end training framework with domain translation and stereo matching networks to tackle this challenge. First, joint optimization between domain translation and stereo matching networks in our end-to-end framework makes the former facilitate the latter one to the maximum extent. Second, this framework introduces two novel losses, i.e., bidirectional multi-scale feature re-projection loss and correlation consistency loss, to help translate all synthetic stereo images into realistic ones as well as maintain epipolar constraints. The effective combination of above two contributions leads to impressive stereo-consistent translation and disparity estimation accuracy. In addition, a mode seeking regularization term is added to endow the synthetic-to-real translation results with higher fine-grained diversity. Extensive experiments demonstrate the effectiveness of the proposed framework on bridging the synthetic-to-real domain gap on stereo matching. Rui Liu 0019, Chengxi Yang, Wenxiu Sun, Xiaogang Wang 0001, Hongsheng Li 0001 |
CVPR | 3 |
| 2020 | Deep Surface Normal Estimation on the 2-Sphere with Confidence Guided Semantic Attention
Quewei Li, Jie Guo 0001, Qinyu Tang, Wenxiu Sun, Jin Zeng 0004, Yanwen Guo 0001 |
ECCV (24) | 5 |
| 2020 | GRNet: Gridding Residual Network for Dense Point Cloud Completion
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Jiageng Mao, Shengping Zhang, Wenxiu Sun |
ECCV (9) | 6 |
| 2020 | Learning Factorized Weight Matrix for Joint FilteringabstractJoint filtering is a fundamental problem in computer vision with applications in many different areas. Most existing algorithms solve this problem with a weighted averaging process to aggregate input pixels. However, the weight matrix of this process is often empirically designed and not robust to complex input. In this work, we propose to learn the weight matrix for joint image filtering. This is a challenging problem, as directly learning a large weight matrix is computationally intractable. To address this issue, we introduce the correlation of deep features to approximate the aggregation weights. However, this strategy only uses inner product for the weight matrix estimation, which limits the performance of the proposed algorithm. Therefore, we further propose to learn a nonlinear function to predict sparse residuals of the feature correlation matrix. Note that the proposed method essentially factorizes the weight matrix into a low-rank and a sparse matrix and then learn both of them simultaneously with deep neural networks. Extensive experiments show that the proposed algorithm compares favorably against the state-of-the-art approaches on a wide variety of joint filtering tasks. Xiangyu Xu 0002, Yongrui Ma, Wenxiu Sun |
ICML | 3 |
| 2020 | Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, Wenxiu Sun |
Int. J. Comput. Vis. | 5 |
| 2020 | Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video DenoisingabstractExisting denoising methods typically restore clear results by aggregating pixels from the noisy input. Instead of relying on hand-crafted aggregation schemes, we propose to explicitly learn this process with deep neural networks. We present a spatial pixel aggregation network and learn the pixel sampling and averaging strategies for image denoising. The proposed model naturally adapts to image structures and can effectively improve the denoised results. Furthermore, we develop a spatio-temporal pixel aggregation network for video denoising to efficiently sample pixels across the spatio-temporal space. Our method is able to solve the misalignment issues caused by large motion in dynamic scenes. In addition, we introduce a new regularization term for effectively training the proposed video denoising model. We present extensive analysis of the proposed method and demonstrate that our model performs favorably against the state-of-the-art image and video denoising approaches on both synthetic and real-world data. Xiangyu Xu 0002, Muchen Li, Wenxiu Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Towards Real Scene Super-Resolution With Raw ImagesabstractMost existing super-resolution methods do not perform well in real scenarios due to lack of realistic training data and information loss of the model input. To solve the first problem, we propose a new pipeline to generate realistic training data by simulating the imaging process of digital cameras. And to remedy the information loss of the input, we develop a dual convolutional neural network to exploit the originally captured radiance information in raw images. In addition, we propose to learn a spatially-variant color transformation which helps more effective color corrections. Extensive experiments demonstrate that super-resolution with raw data helps recover fine details and clear structures, and more importantly, the proposed network and data generation pipeline achieve superior results for single image super-resolution in real scenarios. Xiangyu Xu 0002, Yongrui Ma, Wenxiu Sun |
CVPR | 3 |
| 2019 | Deep Surface Normal Estimation With Hierarchical RGB-D FusionabstractThe growing availability of commodity RGB-D cameras has boosted the applications in the field of scene understanding. However, as a fundamental scene understanding task, surface normal estimation from RGB-D data lacks thorough investigation. In this paper, a hierarchical fusion network with adaptive feature re-weighting is proposed for surface normal estimation from a single RGB-D image. Specifically, the features from color image and depth are successively integrated at multiple scales to ensure global surface smoothness while preserving visually salient details. Meanwhile, the depth features are re-weighted with a confidence map estimated from depth before merging into the color branch to avoid artifacts caused by input depth corruption. Additionally, a hybrid multi-scale loss function is designed to learn accurate normal estimation given noisy ground-truth dataset. Extensive experimental results validate the effectiveness of the fusion strategy and the loss design, outperforming state-of-the-art normal estimation schemes. Jin Zeng 0004, Yanfeng Tong, Yunmu Huang, Qiong Yan, Wenxiu Sun, Jing Chen 0018, Yongtian Wang |
CVPR | 5 |
| 2019 | Deep End-to-End Alignment and Refinement for Time-of-Flight RGB-D ModuleabstractRecently, it is increasingly popular to equip mobile RGB cameras with Time-of-Flight (ToF) sensors for active depth sensing. However, for off-the-shelf ToF sensors, one must tackle two problems in order to obtain high-quality depth with respect to the RGB camera, namely 1) online calibration and alignment; and 2) complicated error correction for ToF depth sensing. In this work, we propose a framework for jointly alignment and refinement via deep learning. First, a cross-modal optical flow between the RGB image and the ToF amplitude image is estimated for alignment. The aligned depth is then refined via an improved kernel predicting network that performs kernel normalization and applies the bias prior to the dynamic convolution. To enrich our data for end-to-end training, we have also synthesized a dataset using tools from computer graphics. Experimental results demonstrate the effectiveness of our approach, achieving state-of-the-art for ToF refinement. Di Qiu, Jiahao Pang, Wenxiu Sun, Chengxi Yang |
ICCV | 3 |
| 2019 | Quadratic Video InterpolationabstractVideo interpolation is an important problem in computer vision, which helps overcome the temporal limitation of camera sensors. Existing video interpolation methods usually assume uniform motion between consecutive frames and use linear models for interpolation, which cannot well approximate the complex motion in the real world. To address these issues, we propose a quadratic video interpolation method which exploits the acceleration information in videos. This method allows prediction with curvilinear trajectory and variable velocity, and generates more accurate interpolation results. For high-quality frame synthesis, we develop a flow reversal layer to estimate flow fields starting from the unknown target frame to the source frame. In addition, we present techniques for flow refinement. Extensive experiments demonstrate that our approach performs favorably against the existing linear models on a wide variety of video datasets. Xiangyu Xu 0002, Li Siyao, Wenxiu Sun, Qian Yin 0001, Ming-Hsuan Yang 0001 |
NeurIPS | 3 |
| 2018 | DSR: Direct Self-Rectification for Uncalibrated Dual-Lens CamerasabstractWith the developments of dual-lens camera modules, depth information representing the third dimension of the captured scenes becomes available for smartphones. It is estimated by stereo matching algorithms, taking as input the two views captured by dual-lens cameras at slightly different viewpoints. Depth-of-field rendering (also be referred to as synthetic defocus or bokeh) is one of the trending depth-based applications. However, to achieve fast depth estimation on smartphones, the stereo pairs need to be rectified in the first place. In this paper, we propose a cost-effective solution to perform stereo rectification for dual-lens cameras called direct self-rectification, short for DSR. It removes the need of individual offline calibration for every pair of dual-lens cameras. In addition, the proposed solution is robust to the slight movements, {\it e.g.}, due to collisions, of the dual-lens cameras after fabrication. Different with existing self-rectification approaches, our approach computes the homography in a novel way with zero geometric distortions introduced to the master image. It is achieved by directly minimizing the vertical displacements of corresponding points between the original master image and the transformed slave image. Our method is evaluated on both realistic and synthetic stereo image pairs, and produces superior results compared to the calibrated rectification or other self-rectification approaches. Ruichao Xiao, Wenxiu Sun, Jiahao Pang, Qiong Yan, Jimmy S. J. Ren |
3DV | 2 |
| 2018 | Single View Stereo MatchingabstractPrevious monocular depth estimation methods take a single view and directly regress the expected results. Though recent advances are made by applying geometrically inspired loss functions during training, the inference procedure does not explicitly impose any geometrical constraint. Therefore these models purely rely on the quality of data and the effectiveness of learning to generalize. This either leads to suboptimal results or the demand of huge amount of expensive ground truth labelled data to generate reasonable results. In this paper, we show for the first time that the monocular depth estimation problem can be reformulated as two sub-problems, a view synthesis procedure followed by stereo matching, with two intriguing properties, namely i) geometrical constraints can be explicitly imposed during inference; ii) demand on labelled depth data can be greatly alleviated. We show that the whole pipeline can still be trained in an end-to-end fashion and this new formulation plays a critical role in advancing the performance. The resulting model outperforms all the previous monocular depth estimation methods as well as the stereo block matching method in the challenging KITTI dataset by only using a small number of real training data. The model also generalizes well to other monocular depth estimation benchmarks. We also discuss the implications and the advantages of solving monocular depth estimation using stereo methods. Jimmy S. J. Ren, Mude Lin, Jiahao Pang, Wenxiu Sun, Hongsheng Li 0001, Liang Lin 0004 |
CVPR | 5 |
| 2018 | LSTM Pose MachinesabstractWe observed that recent state-of-the-art results on single image human pose estimation were achieved by multistage Convolution Neural Networks (CNN). Notwithstanding the superior performance on static images, the application of these models on videos is not only computationally intensive, it also suffers from performance degeneration and flicking. Such suboptimal results are mainly attributed to the inability of imposing sequential geometric consistency, handling severe image quality degradation (e.g. motion blur and occlusion) as well as the inability of capturing the temporal correlation among video frames. In this paper, we proposed a novel recurrent network to tackle these problems. We showed that if we were to impose the weight sharing scheme to the multi-stage CNN, it could be re-written as a Recurrent Neural Network (RNN). This property decouples the relationship among multiple network stages and results in significantly faster speed in invoking the network for videos. It also enables the adoption of Long Short-Term Memory (LSTM) units between video frames. We found such memory augmented RNN is very effective in imposing geometric consistency among frames. It also well handles input quality degradation in videos while successfully stabilizes the sequential outputs. The experiments showed that our approach significantly outperformed current state-of-the-art methods on two large-scale video pose estimation benchmarks. We also explored the memory cells inside the LSTM and provided insights on why such mechanism would benefit the prediction for video-based pose estimations.1 Jimmy S. J. Ren, Zhouxia Wang, Wenxiu Sun, Jinshan Pan, Jiahao Pang, Liang Lin 0004 |
CVPR | 4 |
| 2018 | Zoom and Learn: Generalizing Deep Stereo Matching to Novel DomainsabstractDespite the recent success of stereo matching with convolutional neural networks (CNNs), it remains arduous to generalize a pre-trained deep stereo model to a novel domain. A major difficulty is to collect accurate ground-truth disparities for stereo pairs in the target domain. In this work, we propose a self-adaptation approach for CNN training, utilizing both synthetic training data (with ground-truth disparities) and stereo pairs in the new domain (without ground-truths). Our method is driven by two empirical observations. By feeding real stereo pairs of different domains to stereo models pre-trained with synthetic data, we see that: i) a pre-trained model does not generalize well to the new domain, producing artifacts at boundaries and ill-posed regions; however, ii) feeding an up-sampled stereo pair leads to a disparity map with extra details. To avoid i) while exploiting ii), we formulate an iterative optimization problem with graph Laplacian regularization. At each iteration, the CNN adapts itself better to the new domain: we let the CNN learn its own higher-resolution output; at the meanwhile, a graph Laplacian regularization is imposed to discriminatively keep the desired edges while smoothing out the artifacts. We demonstrate the effectiveness of our method in two domains: daily scenes collected by smart-phone cameras, and street views captured in a driving car. Jiahao Pang, Wenxiu Sun, Chengxi Yang, Jimmy S. J. Ren, Ruichao Xiao, Jin Zeng 0004, Liang Lin 0004 |
CVPR | 2 |
| 2018 | Monocular Depth Estimation with Affinity, Vertical Pooling, and Label Enhancement
Yukang Gan, Xiangyu Xu 0002, Wenxiu Sun, Liang Lin 0004 |
ECCV (3) | 3 |
| 2018 | Convolutional Memory Blocks for Depth Data Representation LearningabstractCompared to natural RGB images, data captured by 3D / depth sensors (e.g., Microsoft Kinect) have different properties, e.g., less discriminable in appearance due to lacking color / texture information. Applying convolutional neural networks (CNNs) on these depth data would lead to unsatisfying learning efficiency, i.e., requiring large amounts of annotated training data for convergence. To address this issue, this paper proposes a novel memory network module, called Convolutional Memory Block (CMB), which empowers CNNs with the memory mechanism on handling depth data. Different from the existing memory networks that store long / short term dependency from sequential data, our proposed CMB focuses on modeling the representative dependency (correlation) among non-sequential samples. Specifically, our CMB consists of one internal memory (i.e., a set of feature maps) and three specific controllers, which enable a powerful yet efficient memory manipulation mechanism. In this way, the internal memory, being implicitly aggregated from all previous inputted samples, can learn to store and utilize representative features among the samples. Furthermore, we employ our CMB to develop a concise framework for predicting articulated pose from still depth images. Comprehensive evaluations on three public benchmarks demonstrate significant superiority (about 6%) of our framework over all the compared methods. More importantly, thanks to the enhanced learning efficiency, our framework can still achieve satisfying results using 50% less training data. Keze Wang, Liang Lin 0004, Chuangjie Ren, Wayne Zhang 0001, Wenxiu Sun |
IJCAI | 5 |
| 2018 | Unsupervised Stereo Matching with Occlusion-Aware Loss
Ningqi Luo, Chengxi Yang, Wenxiu Sun, Binheng Song |
PRICAI (1) | 3 |
| 2017 | Accurate Single Stage Detector Using Recurrent Rolling ConvolutionabstractMost of the recent successful methods in accurate object detection and localization used some variants of R-CNN style two stage Convolutional Neural Networks (CNN) where plausible regions were proposed in the first stage then followed by a second stage for decision refinement. Despite the simplicity of training and the efficiency in deployment, the single stage detection methods have not been as competitive when evaluated in benchmarks consider mAP for high IoU thresholds. In this paper, we proposed a novel single stage end-to-end trainable object detection network to overcome this limitation. We achieved this by introducing Recurrent Rolling Convolution (RRC) architecture over multi-scale feature maps to construct object classifiers and bounding box regressors which are deep in context. We evaluated our method in the challenging KITTI dataset which measures methods under IoU threshold of 0.7. We showed that with RRC, a single reduced VGG-16 based model already significantly outperformed all the previously published results. At the time this paper was written our models ranked the first in KITTI car detection (the hard level), the first in cyclist detection and the second in pedestrian detection. These results were not reached by the previous single stage methods. The code is publicly available. Jimmy S. J. Ren, Xiaohao Chen, Wenxiu Sun, Jiahao Pang, Qiong Yan, Yu-Wing Tai, Li Xu 0001 |
CVPR | 4 |
| 2016 | Look, Listen and Learn - A Multimodal LSTM for Speaker IdentificationabstractSpeaker identification refers to the task of localizing the face of a person who has the same identity as the ongoing voice in a video. This task not only requires collective perception over both visual and auditory signals, the robustness to handle severe quality degradations and unconstrained content variations are also indispensable. In this paper, we describe a novel multimodal Long Short-Term Memory (LSTM) architecture which seamlessly unifies both visual and auditory modalities from the beginning of each sequence input. The key idea is to extend the conventional LSTM by not only sharing weights across time steps, but also sharing weights across modalities. We show that modeling the temporal dependency across face and voice can significantly improve the robustness to content quality degradations and variations. We also found that our multimodal LSTM is robustness to distractors, namely the non-speaking identities. We applied our multimodal LSTM to The Big Bang Theory dataset and showed that our system outperforms the state-of-the-art systems in speaker identification with lower false alarm rate and higher recognition accuracy. Jimmy S. J. Ren, Yongtao Hu 0001, Yu-Wing Tai, Li Xu 0001, Wenxiu Sun, Qiong Yan |
AAAI | 6 |
| 2015 | Shepard Convolutional Neural NetworksabstractDeep learning has recently been introduced to the field of low-level computer vision and image processing. Promising results have been obtained in a number of tasks including super-resolution, inpainting, deconvolution, filtering, etc. However, previously adopted neural network approaches such as convolutional neural networks and sparse auto-encoders are inherently with translation invariant operators. We found this property prevents the deep learning approaches from outperforming the state-of-the-art if the task itself requires translation variant interpolation (TVI). In this paper, we draw on Shepard interpolation and design Shepard Convolutional Neural Networks (ShCNN) which efficiently realizes end-to-end trainable TVI operators in the network. We show that by adding only a few feature maps in the new Shepard layers, the network is able to achieve stronger results than a much deeper architecture. Superior performance on both image inpainting and super-resolution is obtained where our system outperforms previous ones while keeping the running time competitive. Jimmy S. J. Ren, Li Xu 0001, Qiong Yan, Wenxiu Sun |
NIPS | 4 |
| 2015 | Stereo Matching with Optimal Local Adaptive Radiometric CompensationabstractA common assumption in stereo matching is that the corresponding pixels in stereo images have similar pixel values. Unfortunately, such an assumption may not be true due to radiometric variations in different views, leading to severely degraded matching results. In this letter, we propose a radiometrically invariant stereo matching algorithm called Optimal Local Adaptive Radiometric Compensation (LARAC). In LARAC, we approximate the spatially varying Pixel Value Correspondence Function (PVCF) between a corresponding pixel pair as a locally consistent polynomial within an optimal local adaptive window. The optimal polynomial coefficients are obtained for each candidate disparity value and are used to compute the matching cost. Meanwhile, a self-correction property is achieved by the proposed LARAC, leading to reduced matching errors for the outlier pixels. Experimental results suggest that the proposed LARAC outperforms other state-of-the-art stereo matching algorithms. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Feng Zou 0006 |
IEEE Signal Process. Lett. | 3 |
| 2014 | Image compression via sparse reconstructionabstractThe traditional compression system only considers the statistical redundancy of images. Recent compression works exploit the visual redundancy of images to further improve the coding efficiency. However, the existing works only provide suboptimal visual redundancy removal schemes. In this paper, we propose an efficient image compression scheme based on the selection and reconstruction of the visual redundancy. The visual redundancy in an image is defined by some images blocks, named redundant blocks, which can be well reconstructed by the others in the image. At the encoder, we design an effective optimization strategy to elaborately select redundant blocks and intentionally remove them. At the decoder, we propose an image restoration method to reconstruct the removed redundant blocks with minimum reconstructed error. Encouraging experimental results show that our compression scheme achieves up to 13.67% bit rate reduction with a comparable visual quality compared to traditional High Efficiency Video Coding (HEVC). Yuan Yuan 0002, Oscar C. Au, Amin Zheng, Haitao Yang 0001, Ketan Tang, Wenxiu Sun |
ICASSP | 6 |
| 2014 | Natural image matting via adaptive local and nonlocal sample clusteringabstractDigital image matting is the determination of foreground color, background color, and an opacity value of each pixel for an input image. Inherently, matting is a highly ill-posed and under-constrained problem. Thus, some assumptions need to be made to resolve it. Inspired by closed-form matting and color clustering matting, in this work, we first develop an adaptive sample clustering criterion to automatically assign either local or nonlocal neighborhood to each pixel. After that, in order to enhance matting accuracy, we improve the nonlocal clustering performance by introducing a new feature selection parameter to choose preferred feature space for different images in a fully automatic way. And finally we solve the problem using a closed form solution. Experimental results show that our algorithm achieves equal or even better performance among many state-of-the-art matting techniques. Oscar C. Au, Yuan Yuan 0002, Wenxiu Sun, Yonggen Ling, Jiahao Pang |
ICIP | 4 |
| 2014 | Solving dense stereo matching via quadratic programmingabstractWe study the problem of formulating the discrete dense stereo matching using continuous convex optimization. One of the previous work derived a relaxed convex formulation by establishing the relationship between the disparity vector and a warping matrix. However it suffers from high computational complexity. In this paper, the previous convex formulation is translated into an equivalent quadratic programming (QP). Then redundant variables and constraints are eliminated by exploiting the internal sparse property of the warping matrix. The resulting QP can be efficiently tackled using interior point solvers. Moreover, enhanced smoothness term and effective post-processing procedures are also incorporated to further improve the disparity accuracy. Experimental results show that the proposed method is much faster and better than the previous convex formulation, and provides competitive results against existing convex approaches. Oscar C. Au, Pengfei Wan 0001, Wenxiu Sun, Lingfeng Xu, Luheng Jia |
VCIP | 4 |
| 2014 | Seamless View Synthesis Through Texture OptimizationabstractIn this paper, we present a novel view synthesis method named Visto, which uses a reference input view to generate synthesized views in nearby viewpoints. We formulate the problem as a joint optimization of inter-view texture and depth map similarity, a framework that is significantly different from other traditional approaches. As such, Visto tends to implicitly inherit the image characteristics from the reference view without the explicit use of image priors or texture modeling. Visto assumes that each patch is available in both the synthesized and reference views and thus can be applied to the common area between the two views but not the out-of-region area at the border of the synthesized view. Visto uses a Gauss–Seidel-like iterative approach to minimize the energy function. Simulation results suggest that Visto can generate seamless virtual views and outperform other state-of-the-art methods. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Wei Hu 0003 |
IEEE Trans. Image Process. | 1 |
| 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth VideosabstractTransmitting compactly represented geometry of a dynamic 3D scene from a sender can enable a multitude of imaging functionalities at a receiver, such as synthesis of virtual images at freely chosen viewpoints via depth-image-based rendering. While depth maps—projections of 3D geometry onto 2D image planes at chosen camera viewpoints-can nowadays be readily captured by inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. Given depth maps need to be denoised and compressed at the encoder for efficient network transmission to the decoder, in this paper, we consider the denoising and compression problems jointly, arguing that doing so will result in a better overall performance than the alternative of solving the two problems separately in two stages. Specifically, we formulate a rate-constrained estimation problem, where given a set of observed noise-corrupted depth maps, the most probable (maximum a posteriori (MAP)) 3D surface is sought within a search space of surfaces with representation size no larger than a prespecified rate constraint. Our rate-constrained MAP solution reduces to the conventional unconstrained MAP 3D surface reconstruction solution if the rate constraint is loose. To solve our posed rate-constrained estimation problem, we propose an iterative algorithm, where in each iteration the structure (object boundaries) and the texture (surfaces within the object boundaries) of the depth maps are optimized alternately. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint video sequences as input, experimental results show that rate-constrained estimated 3D surfaces computed by our algorithm can reduce coding rate of depth maps by up to 32% compared with unconstrained estimated surfaces for the same quality of synthesized virtual views at the decoder. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 1 |
| 2013 | Ray-space based camera spacing correction via convex optimizationabstract3D technologies such like three-dimensional television and free viewpoint television have caught enormous attentions in the consumer market recently. However, because of the inaccurate camera configuration and environmental constraint, there are errors in the assumed equally-spaced camera intervals. In this paper, we propose a novel camera spacing correction algorithm to detect the spacing errors among the multiple cameras by making the corresponding points co-linear in the epipolar plane images. Experimental results show that the proposed algorithms are robust and can achieve good performance even if the corresponding pixels are not well detected. Meanwhile, our algorithm can be solved by convex optimization with an extremely low complexity. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Wei Hu 0003 |
ICASSP | 3 |
| 2013 | Color clustering mattingabstractNatural image matting refers to the problem of extracting regions of interest such as foreground object from an image based on user inputs like scribbles or trimap. More specifically, we need to estimate the color information of background, foreground and the corresponding opacity, which is an ill-posed problem inherently. Inspired by closed-form matting and KNN matting, in this paper, we extend the local color line model which is based on the assumption of linear color clustering within a small local window, to nonlocal feature space neighborhood. New affinity matrix is defined to achieve better clustering. Further, we demonstrate that good clustering ensures better prediction of alpha matte. Experimental evaluations on benchmark datasets and comparisons show that our matting algorithm is of higher accuracy and better visual quality than some state-of-the-art matting algorithms. Yongfang Shi, Oscar C. Au, Jiahao Pang, Ketan Tang, Wenxiu Sun, Hong Zhang 0024, Luheng Jia |
ICME | 5 |
| 2013 | Rate-distortion optimized 3D reconstruction from noise-corrupted multiview depth videosabstractTransmitting compactly represented geometry of a dynamic scene from a sender can enable a multitude of 3D imaging functionalities at a receiver, such as synthesis of virtual images from freely chosen viewpoints via depth-image-based rendering (DIBR). While depth maps can now be readily captured using inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. In this paper, we derive 3D surfaces of a dynamic scene from noise-corrupted depth maps in a rate-distortion (RD) optimal manner. Specifically, unlike previous work that finds the most likely (e.g., maximum likelihood) 3D surface from noisy observations regardless of representation size, we judiciously search for the best fitting (i.e., minimum distortion) 3D surface subject to a bitrate constraint. Our RD-optimal solution reduces to the maximum likelihood solution as the rate constraint is loosened. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint test sequences as input, experimental results show that RD-optimized 3D reconstructions computed by our algorithm outperform unprocessed depth maps by up to 2:42dB in PSNR of synthesized virtual views at the decoder for the same bitrate. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICME | 1 |
| 2013 | Data hiding in error diffused color halftone imagesabstractHalftone image watermarking has been explored and developed rapidly over the past decade. However, there are still issues to be studied. This paper presents a data hiding method called Data Hiding by Dual Color Conjugate Error Diffusion (DHDCCED) to hide a binary secret pattern into two error diffused color halftone images, such that when the two color halftone images are overlaid, the secret pattern will be revealed. The experimental results show that DHDCCED can significantly improve the performances when comparing both the correct decoding rate and the visual quality of the revealed secret pattern to the existing method Color Conjugate Error Diffusion (CCED). Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang, Wenxiu Sun, Lingfeng Xu |
ISCAS | 5 |
| 2013 | A parallel deblocking filter based on H.264/AVC video coding standardabstractThe deblocking filter in H.264/AVC is one of the most time consuming part of video decoder as its high content adaptation and data dependency lead to lots of computation. In this paper, we propose a novel parallel deblocking filter design based on the H.264/AVC video coding standard, taking the advantage that the data dependency of the deblocking filter are “periodic” in one dimension. Our proposed architecture successfully reduces the dependency between horizontal and vertical filters and utilizes the “periodic” property to achieve pixel-level parallelism. Algorithm analysis and experiment results on JM and GPU show that the proposed deblocking filter keeps as good a coding efficiency as that in H.264/AVC, and its high parallelism is suitable and promising in multi-core/multi-thread computing. Oscar C. Au, Lu Fang 0001, Lin Sun 0004, Wenxiu Sun, Dinuka Soysa |
ISCAS | 5 |
| 2013 | Stereo matching by adaptive weighting selection based cost aggregationabstractCost aggregation is the most essential step for dense stereo correspondence searching, which measures the similarity between pixels in the stereo images. In this paper, based on the analysis of the optimal adaptive weight, we propose a novel support aggregation strategy by adaptive weighting selection. The proposed method calculates the aggregation cost by the joint optimization of both left and right matching cost. By assigning more reasonable weighting coefficients, we exclude the occlusion pixels while preserving sufficient support region for accurate matching. The proposed optimal strategy can be integrated by any other adaptive weighting based cost aggregation method to generate more reasonable similarity measurement. Experimental results show that, compare with traditional methods, our algorithm can reduce the foreground fatten phenomenon while increasing the accuracy in the high texture regions. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Ketan Tang, Yuanfang Guo |
ISCAS | 3 |
| 2013 | Inferring Depth from a Pair of Images Captured Using Different Aperture Settings
Oscar C. Au, Lingfeng Xu, Wenxiu Sun, Wei Hu 0003 |
MMM (2) | 4 |
| 2012 | Novel temporal domain hole filling based on background modeling for view synthesisabstractView synthesis is a technique to generate images/videos in a virtual viewpoint. In this paper, the dis-occlusion/hole problem in view synthesis is resolved from the temporal domain. By the fact that dis-occlusions belong to the background, firstly we build an online background under a newly designed Switchable Gaussian Model (SGM), owning to its computationally simplicity and scene adaptivity. Then, real textures in the dis-occlusions are able to be recovered with the built background. Experimental results have verified the improvements in rendering quality and computation complexity by comparing to the conventional spatial filling methods and other temporal filling methods. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Wei Hu 0003 |
ICIP | 1 |
| 2012 | Adaptive depth map filter for blocking artifacts removal and edge preservingabstractIn depth map coding for 3D video coding systems, coding errors in edges can severely affect the synthesis quality. Edge errors mainly compose of two parts: one is blurring and ringing artifact around sharp edge and the other is fake edge caused by blocking artifact. In this paper, we propose an adaptive depth map filter to remove blocking artifacts while preserving depth edges. The proposed filter is designed based on bilateral filter, in which the range kernel parameter is changed adaptively considering the strength of edges and blocking artifacts. Experimental results demonstrate that the proposed depth map filter can achieve up to 0.41 dB gain on the synthesis quality compared to the deblocking filter in MVC at low bit rate. Wei Hu 0003, Oscar C. Au, Lin Sun 0004, Wenxiu Sun, Lingfeng Xu |
ISCAS | 4 |
| 2012 | Texture optimization for seamless view synthesis through energy minimizationabstractIn this paper, we present a view synthesis method named Visto which aims to generate seamless novel views from a monocular view input. We formulate the problem as joint optimization of inter-view texture similarity and geometry preservation, which significantly differs from traditional view synthesis framework. In this way, the image characteristics of virtual view are inherently inherited from the reference view without introducing any image prior or texture modeling technique. The energy function is minimized using Gauss-Seidel-like approach, and the quality of the virtual view is refined iteratively. The proposed approach also tolerates small depth map errors. Further more, the algorithm is parallel friendly. The simulation results outperform several existing state-of-the-art monocular view synthesis systems. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Wei Hu 0003, Zhiding Yu |
ACM Multimedia | 1 |
| 2011 | Error compensation and reliability based view synthesisabstractView synthesis offers a great flexibility in generating free viewpoint television (FTV) and 3D video (3DV). However, the depth-image-based view synthesis approach is very sensitive to errors in the camera parameters or poorly estimated depth maps (also called depth images). Because of these errors, three kinds of artifacts (blurring, contour, hole) are possibly introduced during the general synthesis process. Comparing to conventional methods which implement the view synthesis only in ideal case, in this paper, we propose to design an error compensation and reliability based view synthesis system where the potential errors are considered. The main contributions are highlighted as follows: Firstly, the camera parameter errors are compensated by a global homography transformation matrix. Secondly, the depth maps are classified into both reliable and unreliable regions and the reliability based weighting masks are built to blend synthesized images from two different views together. Finally, a reliability depth map based hole-filling technique is used to fill the existing holes. The experimental results demonstrate that these artifacts are efficiently reduced in the synthesized images. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Sung Him Chui, Chun Wing Kwok |
ICASSP | 1 |
| 2011 | A convex-optimization approach to dense stereo matchingabstractWe present a novel convex-optimization approach to solving the dense stereo matching problem in computer vision. Instead of directly solving for disparities of pixels, by establishing the connection between a permutation matrix and a disparity vector, we directly formulate the stereo matching problem as a continuous convex quadratic program in a simple, elegant and straightforward manner without performing any complicated relaxation or approximation. By using CVX, the Matlab software for disciplined convex programming, our method is extremely simple to implement. Oscar C. Au, Lingfeng Xu, Wenxiu Sun, Sung Him Chui, Chun Wing Kwok |
ICIP | 4 |
| 2011 | Image rectification for single camera stereo systemabstractSingle camera stereo system utilizes mirrors and a single camera for computational stereo, where the mirrors provide extra views needed for stereo and 3D reconstruction. In this paper, we investigate the basic epiploar geometric properties of the single camera stereo image and propose a novel image rectification technique to map the epipolar lines in the original image into the horizontally aligned lines in the rectified image. Besides, the rotation angles of the single camera corresponding to the planar mirror are derived during rectification. Experimental results show the robustness and accuracy of our method. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Sung Him Chui, Chun Wing Kwok |
ICIP | 3 |
| 2011 | An analysis on bitwise operations in the encrypted domainabstractWhen encrypted data need to be sent to an untrusted computer for processing, it is highly desirable to have a homomorphic cryptosystem that allows the untrusted computer to process the encrypted signals totally without decryption such that, when the encrypted domain result is decrypted, the decrypted value is the same as the equivalent plaintext domain operation. These operations range from basic arithmetic operations to complicated transformations. In this paper, we analyze the existence of equivalent operations in the encrypted domain for four of the common bitwise operations: OR, AND, NOR and NAND in plaintext domain. We will show that such equivalent operations should not exist. For otherwise, if such operations exist, the RSA cryptosystems can be broken by a low-complexity attack. Sung Him Chui, Oscar C. Au, Chun Wing Kwok, Lingfeng Xu, Wenxiu Sun |
ICME | 6 |
| 2011 | Adaptive depth map assisted matting in 3D videoabstractDepth map is widely adopted and available in the 3D research area. Combining the depth map with the matting techniques is helpful to the original matte and depth image based rendering in 3D. Herein, in this paper, a novel adaptive depth map assisted matting approach with concise integration is presented and applied to achieve favorable matting results. In this approach, the Lagrange-multiplier-free closed form solution is firstly derived to reduce the computation complexity and to increase matting accuracy. Based on the work of Levin et al. on closed form matting, an improved alpha matte is then achieved by introducing an adaptive smoothness criterion which is the function of depth map variance. Finally, the matting system is capable of working in a full automatical way by generating the trimap from the depth information. Simulation results demonstrate that the proposed method is able to efficiently generate an alpha matte with an roughly user specified scribbles or an automatically generated trimap. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Zhiding Yu |
ICME | 1 |
| 2011 | Towards robust and efficient segmentation: An approach based on inter-region contour and intra-region content analysisabstractWe address the problem of boundary estimation by formulating it as inter-region contour and intra-region information analysis in the framework of graph-based segmentation. Given an image without any prior information about object model and class, we seek to approximate one's instant perception of visual similarity. The method can serve as a preprocessing step for many higher level operations that require regional support, such as scene understanding and object recognition. We show in this paper that the defined region comparison predicate makes a better boundary estimator than efficient graph-based image segmentation (EGS) - a well known and widely used segmentation method. We further illustrate, by making a small relaxation, further improvement of segmentation performance can be achieved. Experimental results have demonstrated the effectiveness of our proposed method. Zhiding Yu, Oscar C. Au, Ketan Tang, Lingfeng Xu, Wenxiu Sun, Yuanfang Guo |
ICME | 5 |