EDBT 2026 Demo / reviewers in the wild / expert
Jing-Yu Yang 0002
dblp:65/2850-2 · also Jingyu Yang 0002
· DBLP profile ↗
141ranked-venue papers
35as first author
57since 2021 · last 2026
0000-0002-7521-7920ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 110 · 28 first-author · 41 since 2021Artificial intelligence and machine learning · 34 · 4 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained FeaturesabstractThis paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control over fine-grained local editing and free clothing transfer for characters. To enable accurate and intuitive control over the generation and editing of high-quality 3D characters with freely interchangeable clothing, we propose a novel user-interactive approach for disentangled 3D character creation. Specifically, to achieve precise control over 3D character generation and editing, we introduce two user-friendly interaction approaches: a sketch-based layered character generation/editing method, which supports clothing transfer; and a 3D-proxy-based part-level editing method, enabling fine-grained disentangled editing. To enhance 3D character quality, we propose a 3D Gaussian reconstruction strategy guided by geometric priors, ensuring that 3D characters exhibit detailed local geometry and smooth global surfaces. Extensive experiments on both public datasets and in-the-wild data demonstrate that our approach not only generates high-quality disentangled 3D characters but also supports precise and fine-grained editing through user interaction. Zhuo Su 0006, Guidong Wang, Jing-Yu Yang 0002, Yukun Lai, Kun Li 0001 |
AAAI | 5 |
| 2026 | DIC-DDA: Learned Asymmetric Distributed Image Compression via Dual Domain AlignmentabstractMulti-view or stereo image compression is an essential technology in 3D related applications. Due to the overlap between different views, exploring their correlations can help improve the compression rate. However, the computing complexity of joint encoding at the encoding side is a heavy burden for terminal encoders. To solve this problem, the learned Distributed Image Coding (DIC), which only uses the correlated view (namely the side image, SI) in the decoder side, has gained much attention in recent years. In this work, we explore asymmetric DIC where one view is selected as the SI and is losslessly compressed. The key problem in learned asymmetric DIC is alignment between the transmitted low-quality target image and high-quality SI. Previous methods usually adopt patch-level alignment with the offset index obtained from degraded (via re-encoded and decoded) SI and the decoded target image, which hinders the alignment accuracy. In this work, we propose a dual domain alignment strategy, which includes degraded domain and fused domain pixel-wise offset estimation. For the degraded domain alignment, we estimate the offset between the degraded SI feature and the degraded target image feature, which eliminates the difficulties in cross-domain matching. For the fused-domain alignment, we observe that the fusion result of degraded target feature and aligned side image feature implicitly contains fine-scale disparity information. Therefore, we estimate the fine-scale offset from the fusion result, which helps refine the degraded domain offsets. We further propose a selective enhancement module to repair the mismatched region in the aligned feature. Extensive experiments on three datasets demonstrate the superiority of our proposed method, outperforming the second-best method by 16% in terms of average BD-rate reduction on the KITTI Stereo dataset. Our code is available at https://github.com/lixianghuitju/DIC-DDA. Huanjing Yue, Gaosheng Liu, Xin Liu 0012, Jing-Yu Yang 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion ModelabstractDiffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the first framework for zero-shot video restoration and enhancement based on the pre-trained image diffusion model. By replacing the spatial self-attention layer with the proposed short-long-range (SLR) temporal attention layer, the pre-trained image diffusion model can take advantage of the temporal correlation between frames. We further propose temporal consistency guidance, spatial-temporal noise sharing, and an early stopping sampling strategy to improve temporally consistent sampling. Our method is a plug-and-play module that can be inserted into any diffusion-based image restoration or enhancement methods to further improve their performance. Experimental results demonstrate the superiority of our proposed method. Cong Cao 0005, Huanjing Yue, Xin Liu 0012, Jing-Yu Yang 0002 |
AAAI | 4 |
| 2025 | Learning Adaptive Lighting via Channel-Aware GuidanceabstractLearning lighting adaptation is a crucial step in achieving good visual perception and supporting downstream vision tasks. Current research often addresses individual light-related challenges, such as high dynamic range imaging and exposure correction, in isolation. However, we identify shared fundamental properties across these tasks: i) different color channels have different light properties, and ii) the channel differences reflected in the spatial and frequency domains are different. Leveraging these insights, we introduce the channel-aware Learning Adaptive Lighting Network (LALNet), a multi-task framework designed to handle multiple light-related tasks efficiently. Specifically, LALNet incorporates color-separated features that highlight the unique light properties of each color channel, integrated with traditional color-mixed features by Light Guided Attention (LGA). The LGA utilizes color-separated features to guide color-mixed features focusing on channel differences and ensuring visual consistency across all channels. Additionally, LALNet employs dual domain channel modulation for generating color-separated features and a mixed channel modulation and light state space module for producing color-mixed features. Extensive experiments on four representative light-related tasks demonstrate that LALNet significantly outperforms state-of-the-art methods on benchmark tests and requires fewer computational resources. We provide an anonymous online demo at LALNet. Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0026, Huanjing Yue, Jing-Yu Yang 0002 |
ICML | 7 |
| 2025 | FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
Zhuozhao Hu, Kaishen Yuan, Xin Liu 0012, Zitong Yu, Yuan Zong, Jingang Shi, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Multimedia | 8 |
| 2025 | DSDNet: Raw Domain Demoiréing via Dual Color-Space SynergyabstractWith the rapid advancement of mobile imaging, capturing screens using smartphones has become a prevalent practice in distance learning and conference recording. However, moiré artifacts, caused by frequency aliasing between display screens and camera sensors, are further amplified by the image signal processing pipeline, leading to severe visual degradation. Existing sRGB domain demoiréing methods struggle with irreversible information loss, while recent two-stage raw domain approaches suffer from information bottlenecks and inference inefficiency. To address these limitations, we propose a single-stage raw domain demoiréing framework, Dual-Stream Demoiréing Network (DSDNet), which leverages the synergy of raw and YCbCr images to remove moiré while preserving luminance and color fidelity. Specifically, to guide luminance correction and moiré removal, we design a raw-to-YCbCr mapping pipeline and introduce the Synergic Attention with Dynamic Modulation (SADM) module. This module enriches the raw-to-sRGB conversion with cross-domain contextual features. Furthermore, to better guide color fidelity, we develop a Luminance-Chrominance Adaptive Transformer (LCAT), which decouples luminance and chrominance representations. Extensive experiments demonstrate that DSDNet outperforms state-of-the-art methods in both visual quality and quantitative evaluation and achieves an inference speed 2.4x faster than the second-best method, highlighting its practical advantages. We provide an anonymous online demo at https://dsdnet.github.io/DSDNet/. Fangpu Zhang, Yeying Jin, Qihua Cheng, Peng-Tao Jiang, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Multimedia | 7 |
| 2025 | Learning Differential Pyramid Representation for Tone MappingabstractExisting tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. These designs typically fail to preserve fine textures and structural fidelity in complex HDR scenes. Furthermore, most methods lack an effective mechanism to jointly model global tone consistency and local contrast enhancement, leading to globally flat or locally inconsistent outputs such as halo artifacts. We present the Differential Pyramid Representation Network (DPRNet), an end-to-end framework for high-fidelity tone mapping. At its core is a learnable differential pyramid that generalizes traditional Laplacian and Difference-of-Gaussian pyramids through content-aware differencing operations across scales. This allows DPRNet to adaptively capture high-frequency variations under diverse luminance and contrast conditions. To enforce perceptual consistency, DPRNet incorporates global tone perception and local tone tuning modules operating on downsampled inputs, enabling efficient yet expressive tone adaptation. Finally, an iterative detail enhancement module progressively restores the full-resolution output in a coarse-to-fine manner, reinforcing structure and sharpness. Experiments show that DPRNet achieves state-of-the-art results, improving PSNR by **2.39 dB** on the 4K HDR+ dataset and **3.01 dB** on the 4K HDRI Haven dataset, while producing perceptually coherent, detail-preserving results. Demo available at [DPRNet](https://xxxxxxdprnet.github.io/DPRNet/). Yinbo Li, Yihao Liu 0001, Peng-Tao Jiang, Fangpu Zhang, Qihua Cheng, Huanjing Yue, Jing-Yu Yang 0002 |
NeurIPS | 8 |
| 2025 | Multi-Scale Promoted Self-Adjusting Correlation Learning for Facial Action Unit DetectionabstractFacial Action Unit (AU) detection is a crucial task in affective computing and social robotics as it helps to identify emotions expressed through facial expressions. Anatomically, there are innumerable correlations between AUs, which contain rich information and are vital for AU detection. Previous methods used fixed AU correlations based on expert experience or statistical rules on specific benchmarks, but it is challenging to comprehensively reflect complex correlations between AUs via hand-crafted settings. There are alternative methods that employ a fully connected graph to learn these dependencies exhaustively. However, these approaches can result in a computational explosion and high dependency with a large dataset. To address these challenges, this paper proposes a novel self-adjusting AU-correlation learning (SACL) method with less computation for AU detection. This method adaptively learns and updates AU correlation graphs by efficiently leveraging the characteristics of different levels of AU motion and emotion representation information extracted in different stages of the network. Moreover, this paper explores the role of multi-scale learning in correlation information extraction, and design a simple yet effective multi-scale feature learning (MSFL) method to promote better performance in AU detection. By integrating AU correlation information with multi-scale features, the proposed method obtains a more robust feature representation for the final AU detection. Extensive experiments show that the proposed method outperforms the state-of-the-art methods on widely used AU detection benchmark datasets, with only 28.7% and 12.0% of the parameters and FLOPs of the best method, respectively. Xin Liu 0012, Kaishen Yuan, Xuesong Niu, Jingang Shi, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Affect. Comput. | 7 |
| 2025 | RViDeformer: Efficient Raw Video Denoising Transformer With a Larger Benchmark DatasetabstractIn recent years, raw video denoising has garnered increased attention due to the consistency with the imaging process and well-studied noise modeling in the raw domain. However, two problems still hinder the denoising performance. Firstly, there is no large dataset with realistic motions for supervised raw video denoising, as capturing noisy and clean frames for real dynamic scenes is difficult. To address this, we propose recapturing existing high-resolution videos displayed on a 4K screen with high-low ISO settings to construct noisy-clean paired frames. In this way, we construct a video denoising dataset (named as ReCRVD) with 120 groups of noisy-clean videos, whose ISO values ranging from 1600 to 25600. Secondly, while non-local temporal-spatial attention is beneficial for denoising, it often leads to heavy computation costs. We propose an efficient raw video denoising transformer network (RViDeformer) that explores both short and long-distance correlations. Specifically, we propose multi-branch spatial and temporal attention modules, which explore the patch correlations from local window, local low-resolution window, global downsampled window, and neighbor-involved window, and then they are fused together. We employ reparameterization to reduce computation costs. Our network is trained in both supervised and unsupervised manners, achieving the best performance compared with state-of-the-art methods. Additionally, the model trained with our proposed dataset (ReCRVD) outperforms the model trained with previous benchmark dataset (CRVD) when evaluated on the real-world outdoor noisy videos.Our code and dataset will be released. Huanjing Yue, Cong Cao 0005, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | TransDiff: Unsupervised Non-Line-of-Sight Imaging With Aperture-Limited Relay SurfacesabstractNon-line-of-sight (NLOS) imaging aims to reconstruct scenes hidden from direct view and has broad applications in robotic vision, rescue operations, autonomous driving, and remote sensing. However, most existing methods rely on densely sampled transients from large, continuous relay surfaces, which limits their practicality in real-world scenarios with aperture constraints. To address this limitation, we propose an unsupervised zero-shot framework tailored for confocal NLOS imaging with aperture-limited relay surfaces. Our method leverages latent diffusion models to recover fully-sampled transients from undersampled versions by enforcing measurement consistency during the sampling process. To further improve recovered transient quality, we introduce a progressive recovery strategy that incrementally recovers missing transient values, effectively mitigating the impact of severe aperture limitations. In addition, to suppress error propagation during recovery, we develop a backpropagation-based error correction reconstruction algorithm that refines intermediate recovered transients by enforcing sparsity regularization in the voxel domain, enabling high-fidelity final reconstructions. Extensive experiments on both simulated and real-world datasets validate the robustness and generalization capability of our method across diverse aperture-limited relay surfaces. Notably, our method follows a zero-shot paradigm, requiring only a single pretraining stage without paired data or pattern-specific retraining, which makes it a more practical and generalizable framework for NLOS imaging. Xingyu Cui, Huanjing Yue, Shida Sun, Yusen Hou, Zhiwei Xiong, Jing-Yu Yang 0002 |
IEEE Trans. Image Process. | 7 |
| 2025 | A Simple Yet Effective Network Based on Vision Transformer for Camouflaged Object and Salient Object DetectionabstractCamouflaged object detection (COD) and salient object detection (SOD) are two distinct yet closely-related computer vision tasks widely studied during the past decades. Though sharing the same purpose of segmenting an image into binary foreground and background regions, their distinction lies in the fact that COD focuses on concealed objects hidden in the image, while SOD concentrates on the most prominent objects in the image. Building universal segmentation models is currently a hot topic in the community. Previous works achieved good performance on certain task by stacking various hand-designed modules and multi-scale features. However, these careful task-specific designs also make them lose their potential as general-purpose architectures. Therefore, we hope to build general architectures that can be applied to both tasks. In this work, we propose a simple yet effective network (SENet) based on vision Transformer (ViT), by employing a simple design of an asymmetric ViT-based encoder-decoder structure, we yield competitive results on both tasks, exhibiting greater versatility than meticulously crafted ones. To enhance the performance of universal architectures on both tasks, we propose some general methods targeting some common difficulties of the two tasks. First, we use image reconstruction as an auxiliary task during training to increase the difficulty of training, forcing the network to have a better perception of the image as a whole to help with segmentation tasks. In addition, we propose a local information capture module (LICM) to make up for the limitations of the patch-level attention mechanism in pixel-level COD and SOD tasks and a dynamic weighted loss (DW loss) to solve the problem that small target samples are more difficult to locate and segment in both tasks. Finally, we also conduct a preliminary exploration of joint training, trying to use one model to complete two tasks simultaneously. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our method. The code is available at https://github.com/linuxsino/SENet. Chao Hao, Zitong Yu, Xin Liu 0012, Jun Xu 0019, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | Learning to See Low-Light Images via Feature Domain AdaptationabstractRaw low-light image enhancement (LLIE) has achieved much better performance than the sRGB domain enhancement methods due to the merits of raw data. However, the ambiguity between noisy to clean and raw to sRGB mappings may mislead the single-stage enhancement networks. The two-stage networks avoid ambiguity by step-by-step or decoupling the two mappings but usually have large computing complexity. To solve this problem, we propose a single-stage network empowered by Feature Domain Adaptation (FDA) to decouple the denoising and color mapping tasks in raw LLIE. The denoising encoder is supervised by the clean raw image, and then the denoised features are adapted for the color mapping task by an FDA module. We propose a Lineformer to serve as the FDA, which can well explore the global and local correlations with fewer line buffers (friendly to the line-based imaging process). During inference, the raw supervision branch is removed. In this way, our network combines the advantage of a two-stage enhancement process with the efficiency of single-stage inference. Experiments on four benchmark datasets demonstrate that our method achieves state-of-the-art performance with fewer computing costs (60% FLOPs of the two-stage method DNF). Our codes will be released after the acceptance of this work. Qihua Cheng, Huanjing Yue, Yihao Liu 0001, Jing-Yu Yang 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | JASRNet: Learning Joint Adaptive Sampling and Reconstruction for Depth SensingabstractRecent attempts to exploit irregular sampling strategies for depth sensing have shown prominent merits over the uniform rectangular sampling in terms of depth reconstruction quality, particularly at low sampling rates. However, the separate treatment of depth sampling and reconstruction did not enjoy potential merits of joint optimization. In this article, we propose a joint adaptive depth sampling and reconstruction network, named JASRNet , for the RGB-D sensing configuration, to simultaneously optimize both the sampling and reconstruction of the depth information in an end-to-end manner. The sampling sub-network infers the locations to sample according to the significance distribution generated from the associated RGB image without any prior information of the underlying depth maps. The depth reconstruction sub-network learns and then fuses global and local depth features with attention guidance, which helps to obtain more accurate depth reconstruction results at boundaries. A hybrid loss function is further proposed to promote sharp discontinuities of the reconstructed depth maps. The qualitative and quantitative results show that our method achieves better depth sensing quality than several state-of-the-art methods for various indoor and outdoor scenes. Chunyang Bi, Mingnuo Teng, Tianhao Xie, Kun Li 0001, Jing-Yu Yang 0002 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2025 | Learned Focused Plenoptic Image Compression With Local-Global Correlation LearningabstractThe dense light field sampling of focused plenoptic images (FPIs) yields substantial amounts of redundant data, necessitating efficient compression in practical applications. However, the presence of discontinuous structures and long-distance properties in FPIs poses a challenge. In this paper, we propose a novel end-to-end approach for learned focused plenoptic image compression (LFPIC). Specifically, we introduce a local-global correlation learning strategy to build the nonlinear transforms. This strategy can effectively handle the discontinuous structures and leverage long-distance correlations in FPI for high compression efficiency. Additionally, we propose a spatial-wise context model tailored for LFPIC to help emphasize the most related symbols during coding and further enhance the rate-distortion performance. Experimental results demonstrate the effectiveness of our proposed method, achieving a 22.16% BD-rate reduction (measured in PSNR) on the public dataset compared to the recent state-of-the-art LFPIC method. This improvement holds significant promise for benefiting the applications of focused plenoptic cameras. Gaosheng Liu, Huanjing Yue, Bihan Wen, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | High-Quality Reconstruction of Depth Maps From Graph-Based Non-Uniform SamplingabstractDepth sensing is essential for intelligent computer vision applications, but it often suffers from low range precision and spatial resolution. To address this problem, we propose a novel framework that combines non-uniform sampling and reconstruction based on graph theory. Our framework consists of two main components: (1) a graph Laplacian induced non-uniform sampling (GLINUS) scheme that samples depth signals more densely around edges and contours than in smooth regions, and (2) an ensemble of priors (EoP) model that reconstructs the high-quality depth map using adaptive dual-tree discrete wavelet packets (ADDWP) transform, graph total variation regularizer, and graph Laplacian regularizer with color guidance. We solve the reconstruction problem using the alternating direction method of multipliers (ADMM). Our experiments demonstrate that our framework can capture fine structures and global information in depth signals and produce superior depth reconstruction results. Jing-Yu Yang 0002, Yusen Hou, Xinchen Ye, Pascal Frossard, Kun Li 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free MatchingabstractDual-lens super-resolution (SR) is a practical scenario for reference (Ref) based SR by utilizing the telephoto image (Ref) to assist the super-resolution of the low-resolution wide-angle image (LR input). Different from general RefSR, the Ref in dual-lens SR only covers the overlapped field of view (FoV) area. However, current dual-lens SR methods rarely utilize these specific characteristics and directly perform dense matching between the LR input and Ref. Due to the resolution gap between LR and Ref, the matching may miss the best-matched candidate and destroy the consistent structures in the overlapped FoV area. Different from them, we propose to first align the Ref with the center region (namely the overlapped FoV area) of the LR input by combining global warping and local warping to make the aligned Ref be sharp and consistent. Then, we formulate the aligned Ref and LR center as value-key pairs, and the corner region of the LR is formulated as queries. In this way, we propose a kernel-free matching strategy by matching between the LR-corner (query) and LR-center (key) regions, and the corresponding aligned Ref (value) can be warped to the corner region of the target. Our kernel-free matching strategy avoids the resolution gap between LR and Ref, which makes our network have better generalization ability. In addition, we construct a DuSR-Real dataset with (LR, Ref, HR) triples, where the LR and HR are well aligned. Experiments on three datasets demonstrate that our method outperforms the second-best method by a large margin. Our code and dataset are available at https://github.com/ZifanCui/KeDuSR. Huanjing Yue, Zifan Cui, Kun Li 0001, Jing-Yu Yang 0002 |
AAAI | 4 |
| 2024 | LPSNet: End-to-End Human Pose and Shape Estimation with Lensless ImagingabstractHuman pose and shape (HPS) estimation with lensless imaging is not only beneficial to privacy protection but also can be used in covert surveillance scenarios due to the small size and simple structure of this device. However, this task presents significant challenges due to the inherent ambiguity of the captured measurements and lacks effective methods for directly estimating human pose and shape from lensless data. In this paper, we propose the first end-to-end framework to recover 3D human poses and shapes from lensless measurements to our knowledge. We specifically design a multi-scale lensless feature decoder to decode the lensless measurements through the optically encoded mask for efficient feature extraction. We also propose a double-head auxiliary supervision mechanism to improve the estimation accuracy of human limb ends. Besides, we establish a lensless imaging system and verify the effectiveness of our method on various datasets acquired by our lensless imaging system. The code and dataset are available at https://cic.tju.edu.cn/faculty/likun/projects/LPSNet. Haoyang Ge, Qiao Feng 0001, Hailong Jia, Xiongzheng Li, Xiangjun Yin, Jing-Yu Yang 0002, Kun Li 0001 |
CVPR | 7 |
| 2024 | AUFormer: Vision Transformers Are Parameter-Efficient Facial Action Unit Detectors
Kaishen Yuan, Zitong Yu, Xin Liu 0012, Weicheng Xie 0001, Huanjing Yue, Jing-Yu Yang 0002 |
ECCV (50) | 6 |
| 2024 | Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality ReweighterabstractDeep neural networks (DNNs) have been applied in many computer vision tasks and achieved state-of-the-art (SOTA) performance. However, misclassification will occur when DNNs predict adversarial examples which are created by adding human-imperceptible adversarial noise to natural examples. This limits the application of DNN in security-critical fields. In order to enhance the robustness of models, previous research has primarily focused on the unimodal domain, such as image recognition and video understanding. Although multi-modal learning has achieved advanced performance in various tasks, such as action recognition, research on the robustness of RGB-skeleton action recognition models is scarce. In this paper, we systematically investigate how to improve the robustness of RGB-skeleton action recognition models. We initially conducted empirical analysis on the robustness of different modalities and observed that the skeleton modality is more robust than the RGB modality. Motivated by this observation, we propose the Attention-based Modality Reweighter (AMR), which utilizes an attention layer to re-weight the two modalities, enabling the model to learn more robust features. Our AMR is plug-and-play, allowing easy integration with multimodal models. To demonstrate the effectiveness of AMR, we conducted extensive experiments on various datasets. For example, compared to the SOTA methods, AMR exhibits a 43.77% improvement against PGD20 attacks on the NTURGB+D 60 dataset. Furthermore, it effectively balances the differences in robustness between different modalities. Xin Liu 0012, Zitong Yu, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002 |
IJCB | 6 |
| 2024 | Efficient Screen Content Image Compression via Superpixel-based Content Aggregation and Dynamic Feature Fusion
Sheng Shen 0010, Huanjing Yue, Jing-Yu Yang 0002 |
IJCAI | 3 |
| 2024 | Virtual Scanning: Unsupervised Non-line-of-sight Imaging from Irregularly Undersampled TransientsabstractNon-line-of-sight (NLOS) imaging allows for seeing hidden scenes around corners through active sensing.
Most previous algorithms for NLOS reconstruction require dense transients acquired through regular scans over a large relay surface, which limits their applicability in realistic scenarios with irregular relay surfaces.
In this paper, we propose an unsupervised learning-based framework for NLOS imaging from irregularly undersampled transients~(IUT).
Our method learns implicit priors from noisy irregularly undersampled transients without requiring paired data, which is difficult and expensive to acquire and align.
To overcome the ambiguity of the measurement consistency constraint in inferring the albedo volume, we design a virtual scanning process that enables the network to learn within both range and null spaces for high-quality reconstruction.
We devise a physics-guided SURE-based denoiser to enhance robustness to ubiquitous noise in low-photon imaging conditions.
Extensive experiments on both simulated and real-world data validate the performance and generalization of our method.
Compared with the state-of-the-art (SOTA) method, our method achieves higher fidelity, greater robustness, and remarkably faster inference times by orders of magnitude.
The code and model are available at https://github.com/XingyuCuii/Virtual-Scanning-NLOS. Xingyu Cui, Huanjing Yue, Xiangjun Yin, Yusen Hou, Yun Meng, Jing-Yu Yang 0002 |
NeurIPS | 9 |
| 2024 | Unsupervised HDR Image and Video Tone Mapping via Contrastive LearningabstractCapturing high dynamic range (HDR) images (videos) is attractive because it can reveal the details in both dark and bright regions. Since the mainstream screens only support low dynamic range (LDR) content, tone mapping algorithm is required to compress the dynamic range of HDR images (videos). Although image tone mapping has been widely explored, video tone mapping is lagging behind, especially for the deep-learning-based methods, due to the lack of HDR-LDR video pairs. In this work, we propose a unified framework (IVTMNet) for unsupervised image and video tone mapping. To improve unsupervised training, we propose domain and instance based contrastive learning loss. Instead of using a universal feature extractor, such as VGG to extract the features for similarity measurement, we propose a novel latent code, which is an aggregation of the brightness and contrast of extracted features, to measure the similarity of different pairs. We totally construct two negative pairs and three positive pairs to constrain the latent codes of tone mapped results. For the network structure, we propose a spatial-feature-enhanced (SFE) module to enable information exchange and transformation of nonlocal regions. For video tone mapping, we propose a temporal-feature-replaced (TFR) module to efficiently utilize the temporal correlation and improve the temporal consistency of video tone-mapped results. We construct a large-scale unpaired HDR-LDR video dataset to facilitate the unsupervised training process for video tone mapping. Experimental results demonstrate that our method outperforms state-of-the-art image and video tone mapping methods. Our code and dataset are available athttps://github.com/cao-cong/UnCLTMO. Cong Cao 0005, Huanjing Yue, Xin Liu 0012, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | High-Quality Animatable Dynamic Garment Reconstruction From Monocular VideosabstractMuch progress has been made in reconstructing garments from an image or a video. However, none of existing works meet the expectations of digitizing high-quality animatable dynamic garments that can be adjusted to various unseen poses. In this paper, we propose the first method to recover high-quality animatable dynamic garments from monocular videos without depending on scanned data. To generate reasonable deformations for various unseen poses, we propose a learnable garment deformation network that formulates the garment reconstruction task as a pose-driven deformation problem. To alleviate the ambiguity estimating 3D garments from monocular videos, we design a multi-hypothesis deformation module that learns spatial representations of multiple plausible deformations. Experimental results on several public datasets demonstrate that our method can reconstruct high-quality dynamic garments with coherent surface details, which can be easily animated under unseen poses. The code will be provided for research purposes. Xiongzheng Li, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | rPPG-MAE: Self-Supervised Pretraining With Masked Autoencoders for Remote Physiological MeasurementsabstractRemote photoplethysmography (rPPG) is an important technique for detecting human vital signs and has received extensive attention. For a long time, researchers have focused attention on supervised methods that rely on large amounts of labeled data. These methods are limited by their need for large amounts of data and the difficulty of acquiring ground truth physiological signals. To address these issues, several self-supervised methods based on contrastive learning have been proposed. However, they focus on contrastive learning between samples, which neglects inherent self-similar priors in physiological signals and seems to have a limited ability to cope with noise. In this paper, a linear self-supervised reconstruction task was designed for extracting the inherent self-similar priors in physiological signals. In addition, a specific noise-insensitive strategy was explored for reducing the interference of motion and illumination. The framework proposed in this paper, rPPG-MAE, demonstrates excellent performance even on the challenging VIPL-HR dataset. We also evaluate the proposed method on two public datasets, namely, PURE and UBFC-rPPG. The results show that our method not only outperforms existing self-supervised methods but also outperforms state-of-the-art (SOTA) supervised methods. One important observation is that the quality of the dataset appears to be more important than the size of the dataset used in self-supervised pretraining of the rPPG. The source code is available athttps://github.com/linuxsino/rPPG-MAE. Xin Liu 0012, Yuting Zhang 0008, Zitong Yu, Hao Lu 0009, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 6 |
| 2024 | From Recognition to Prediction: Leveraging Sequence Reasoning for Action AnticipationabstractThe action anticipation task refers to predicting what action will happen based on observed videos, which requires the model to have a strong ability to summarize the present and then reason about the future. Experience and common sense suggest that there is a significant correlation between different actions, which provides valuable prior knowledge for the action anticipation task. However, previous methods have not effectively modeled this underlying statistical relationship. To address this issue, we propose a novel end-to-end video modeling architecture that utilizes attention mechanisms, named Anticipation via Recognition and Reasoning (ARR). ARR decomposes the action anticipation task into action recognition and sequence reasoning tasks and effectively learns the statistical relationship between actions by next action prediction (NAP). In comparison to existing temporal aggregation strategies, ARR is able to extract more effective features from observable videos to make more reasonable predictions. In addition, to address the challenge of relationship modeling that requires extensive training data, we propose an innovative approach for the unsupervised pre-training of the decoder, which leverages the inherent temporal dynamics of video to enhance the reasoning capabilities of the network. Extensive experiments on the Epic-kitchen-100, EGTEA Gaze+, and 50salads datasets demonstrate the efficacy of the proposed methods. The code is available at https://github.com/linuxsino/ARR . Xin Liu 0012, Chao Hao, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Learning Semantic-Aware Disentangled Representation for Flexible 3D Human Body Editingabstract3D human body representation learning has received increasing attention in recent years. However, existing works cannot flexibly, controllably and accurately represent human bodies, limited by coarse semantics and unsatisfactory representation capability, particularly in the absence of supervised data. In this paper, we propose a human body representation with fine-grained semantics and high reconstruction-accuracy in an unsupervised setting. Specifically, we establish a correspondence between latent vectors and geometric measures of body parts by designing a part-aware skeleton-separated decoupling strategy, which facilitates controllable editing of human bodies by modifying the corresponding latent codes. With the help of a bone-guided auto-encoder and an orientation-adaptive weighting strategy, our representation can be trained in an unsupervised manner. With the geometrically meaningful latent space, it can be applied to a wide range of applications, from human body editing to latent code interpolation and shape style transfer. Experimental results on public datasets demonstrate the accurate reconstruction and flexible editing abilities of the proposed method. The code will be available at http://cic.tju.edu.cn/faculty/likun/projects/SemanticHuman. Xiaokun Sun, Qiao Feng 0001, Xiongzheng Li, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
CVPR | 6 |
| 2023 | Dec-Adapter: Exploring Efficient Decoder-Side Adapter for Bridging Screen Content and Natural Image CompressionabstractNatural image compression has been greatly improved in the deep learning era. However, the compression performance will be heavily degraded if the pretrained encoder is directly applied on screen content image compression. Meanwhile, we observe that parameter-efficient transfer learning (PETL) methods have shown great adaptation ability in high-level vision tasks. Therefore, we propose a Dec-Adapter, a pioneering entropy-efficient transfer learning module for the decoder to bridge natural image and screen content compression. The adapter’s parameters are learned during encoding and transmitted to the decoder for image-adaptive decoding. Our Dec-Adapter is lightweight, domain-transferable, and architecture-agnostic with generalized performance in bridging the two domains. Experiments demonstrate that our method outperforms all existing methods by a large margin in terms of BD-rate performance on screen content image compression. Specifically, our method achieves over 2 dB gain compared with the baseline when transferred to screen content image compression. Sheng Shen 0010, Huanjing Yue, Jing-Yu Yang 0002 |
ICCV | 3 |
| 2023 | NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigationabstract3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general contains much more information than 2D image; (ii) many 3D representations are not well suited for gradient-based optimization, let alone disentanglement. To address these challenges, we use NeRF as a differentiable 3D representation, and introduce a self-supervised Navigation to identify interpretable semantic directions in the latent space. To our best knowledge, this novel method, dubbed NaviNeRF, is the first work to achieve fine-grained 3D disentanglement without any priors or supervisions. Specifically, NaviNeRF is built upon the generative NeRF pipeline, and equipped with an Outer Navigation Branch and an Inner Refinement Branch. They are complementary —— the outer navigation is to identify global-view semantic directions, and the inner refinement dedicates to fine-grained attributes. A synergistic loss is further devised to coordinate two branches. Extensive experiments demonstrate that NaviNeRF has a superior fine-grained 3D disentanglement ability than the previous 3D-aware models. Its performance is also comparable to editing-oriented models relying on semantic or geometry priors.* Baao Xie, Bohan Li 0015, Zequn Zhang, Junting Dong, Xin Jin 0014, Jing-Yu Yang 0002, Wenjun Zeng 0001 |
ICCV | 6 |
| 2023 | Recaptured Raw Screen Image and Video Demoiréing via Channel and Spatial ModulationsabstractCapturing screen contents by smartphone cameras has become a common way for information sharing. However, these images and videos are often degraded by moiré patterns, which are caused by frequency aliasing between the camera filter array and digital display grids. We observe that the moiré patterns in raw domain is simpler than those in sRGB domain, and the moiré patterns in raw color channels have different properties. Therefore, we propose an image and video demoiréing network tailored for raw inputs. We introduce a color-separated feature branch, and it is fused with the traditional feature-mixed branch via channel and spatial modulations. Specifically, the channel modulation utilizes modulated color-separated features to enhance the color-mixed features. The spatial modulation utilizes the feature with large receptive field to modulate the feature with small receptive field. In addition, we build the first well-aligned raw video demoiréing (RawVDemoiré) dataset and propose an efficient temporal alignment method by inserting alternating patterns. Experiments demonstrate that our method achieves state-of-the-art performance for both image and video demoiréing. Our dataset and code will be released after the acceptance of this work. Yijia Cheng, Xin Liu 0012, Jing-Yu Yang 0002 |
NeurIPS | 3 |
| 2023 | Adversarial Dual-Student With Differentiable Spatial Warping for Semi-Supervised Semantic SegmentationabstractA common challenge posed to robust semantic segmentation is the expensive data annotation cost. Existing semi-supervised solutions show great potential for solving this problem. Their key idea is constructing consistency regularization with unsupervised data augmentation from unlabeled data for model training. The perturbations for unlabeled data enable the consistency training loss, which benefits semi-supervised semantic segmentation. However, these perturbations destroy image context and introduce unnatural boundaries, which is harmful for semantic segmentation. Besides, the widely adopted semi-supervised learning framework, i.e. mean-teacher, suffers performance limitation since the student model finally converges to the teacher model. In this paper, first of all, we propose a context friendly differentiable geometric warping to conduct unsupervised data augmentation; secondly, a novel adversarial dual-student framework is proposed to improve the Mean-Teacher from the following two aspects: (1) dual student models are learned independently except for a stabilization constraint to encourage exploiting model diversities; (2) adversarial training scheme is applied to both students and the discriminators are resorted to distinguish reliable pseudo-label of unlabeled data for self-training. Effectiveness is validated via extensive experiments on PASCAL VOC2012 and Cityscapes. Our solution significantly improves the performance and state-of-the-art results are achieved on both datasets. Remarkably, compared with fully supervision, our solution achieves comparable mIoU of 73.4% using only 12.5% annotated data on PASCAL VOC2012. Our codes and models are available athttps://github.com/cao-cong/ADS-SemiSeg. Cong Cao 0005, Dongliang He, Fu Li 0003, Huanjing Yue, Jing-Yu Yang 0002, Errui Ding |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Intra-Inter View Interaction Network for Light Field Image Super-ResolutionabstractLight field (LF) cameras, which can record real-word scenes from multiple viewpoints in a single shot, are widely used in 3D reconstruction, re-focusing, and virtual realityetc. However, the inherent trade-off between spatial resolution and angular resolution of LF images hinders their applications for scenarios requiring high resolutions. In this paper, we propose a novel intra-inter view interaction network for LF image super-resolution, termed as LF-IINet, to exploit the correlations among all views and simultaneously preserve the parallax structure of LF views. The proposed LF-IINet consists of two parallel branches. Specifically, the top branch extracts global inter-view information, and the bottom branch first independently maps each view to deep representations and then models the correlations among all intra-view features via proposed multi-view context block (MCB). The two branches interact with each other by proposed inter-assist-intra feature updating module (IntraFUM, where the intra feature are updated with the assistance of the inter feature) and intra-assist-inter feature updating module (InterFUM, where the inter feature are updated with the assistance of the intra feature). In this way, our LF-IINet incorporates rich angular and spatial information for LF image super-resolution. Extensive comparison with state-of-the-art methods demonstrates that our method achieves superior performance visually and quantitatively. Furthermore, quantitative results also show that our method is effective for LF images with either small or large disparities. Our code is shared inhttps://github.com/GaoshengLiu/LF-IINet. Gaosheng Liu, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 4 |
| 2023 | Efficient Light Field Angular Super-Resolution With Sub-Aperture Feature Learning and Macro-Pixel UpsamplingabstractThe acquisition of densely-sampled light field (LF) images is costly, which hampers the applications of LF imaging technology in 3D reconstruction, digital refocusing, virtual reality,etc. To mitigate the obstacle, various approaches have been proposed to reconstruct densely-sampled LF images from sparsely-sampled ones. However, most existing methods still suffer from the non-Lambertian effect and large disparity issue. In this paper, we embrace the challenges by introducing a new paradigm for LF angular super-resolution (SR), which first explores the multi-scale spatial-angular correlations on the sparse sub-aperture images (SAIs) and then performs angular SR on macro-pixel features. In this way, we propose an efficient LF angular SR network, termed as EASR, with simple 3D (2D) CNNs and reshaping operations. The proposed EASR can extract effective feature representations on SAIs and can handle large disparities well by performing angular SR on macro-pixel features. Extensive comparisons with state-of-the-art methods demonstrate that our method achieves superior performance visually and quantitatively. Furthermore, our method achieves efficient angular SR by providing an excellent tradeoff between reconstruction performance and inference time. Gaosheng Liu, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 4 |
| 2023 | Recaptured Screen Image Demoiréing in Raw DomainabstractCapturing screen content by smart-phone cameras has become a daily routine to record or share instant information from display screens for convenience. However, these recaptured screen images are often degraded by moiré patterns and usually present color cast against the original screen source. We observe that performing demoiréing in raw domain before feeding into the image signal processor (ISP) is more effective than demoiréing in the sRGB domain as done in recent demoiréing works. In this paper, we investigate the demoiréing of raw images through a class-specific learning approach. To this end, we build the first well-aligned raw moiré image dataset by pixel-wise alignment between the recaptured images and source ones. Noting that document images occupy a large portion of screen contents and have different properties from generic images, we propose a class-specific learning strategy for textual images and natural color images. In addition, to deal with moiré patterns with various scales, a multi-scale encoder with multi-level feature fusion is proposed. The shared encoder enables us to extract rich representations for the two kinds of contents and the class-specific decoders benefit the specific content reconstruction by focusing on targeted representations. Experiment results demonstrate that our method achieves state-of-the-art demoiréing performance. We have released the code and dataset inhttps://github.com/tju-chengyijia/RDNet Huanjing Yue, Yijia Cheng, Cong Cao 0005, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 5 |
| 2023 | Learning to Infer Inner-Body Under Clothing From Monocular VideoabstractAccurately estimating the human inner-body under clothing is very important for body measurement, virtual try-on and VR/AR applications. In this article, we propose the first method to allow everyone to easily reconstruct their own 3D inner-body under daily clothing from a self-captured video with the mean reconstruction error of 0.73cm within 15s. This avoids privacy concerns arising from nudity or minimal clothing. Specifically, we propose a novel two-stage framework with a Semantic-guided Undressing Network (SUNet) and an Intra-Inter Transformer Network (IITNet). SUNet learns semantically related body features to alleviate the complexity and uncertainty of directly estimating 3D inner-bodies under clothing. IITNet reconstructs the 3D inner-body model by making full use of intra-frame and inter-frame information, which addresses the misalignment of inconsistent poses in different frames. Experimental results on both public datasets and our collected dataset demonstrate the effectiveness of the proposed method. The code and dataset is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/Inner-Body. Xiongzheng Li, Xiaokun Sun, Haibiao Xuan, Yukun Lai, Yingdi Xie, Jing-Yu Yang 0002, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2022 | Reference-Based Speech Enhancement via Feature Alignment and Fusion NetworkabstractSpeech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference speech) as a global embedding to guide the enhancement process. Different from them, we observe that the speeches of the same speaker are correlated in terms of frame-level short-time Fourier Transform (STFT) spectrogram. Therefore, we propose reference-based speech enhancement via a feature alignment and fusion network (FAF-Net). Given a noisy speech and a clean reference speech spoken by the same speaker, we first propose a feature level alignment strategy to warp the clean reference with the noisy speech in frame level. Then, we fuse the reference feature with the noisy feature via a similarity-based fusion strategy. Finally, the fused features are skipped connected to the decoder, which generates the enhanced results. Experimental results demonstrate that the performance of the proposed FAF-Net is close to state-of-the-art speech enhancement methods on both DNS and Voice Bank+DEMAND datasets. Our code is available at https://github.com/HieDean/FAF-Net. Huanjing Yue, Wenxin Duo, Xiulian Peng, Jing-Yu Yang 0002 |
AAAI | 4 |
| 2022 | Real-RawVSR: Real-World Raw Video Super-Resolution with a Benchmark Dataset
Huanjing Yue, Jing-Yu Yang 0002 |
ECCV (6) | 3 |
| 2022 | FOF: Learning Fourier Occupancy Field for Monocular Real-time Human ReconstructionabstractThe advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties achieving high-quality results and real-time speed at the same time. In this paper, we propose Fourier Occupancy Field (FOF), a novel, powerful, efficient and flexible 3D geometry representation, for monocular real-time and accurate human reconstruction. A FOF represents a 3D object with a 2D field orthogonal to the view direction where at each 2D position the occupancy field of the object along the view direction is compactly represented with the first few terms of Fourier series, which retains the topology and neighborhood relation in the 2D domain. A FOF can be stored as a multi-channel image, which is compatible with 2D convolutional neural networks and can bridge the gap between 3D geometries and 2D images. A FOF is very flexible and extensible, \eg, parametric models can be easily integrated into a FOF as a prior to generate more robust results. Meshes and our FOF can be easily inter-converted. Based on FOF, we design the first 30+FPS high-fidelity real-time monocular human reconstruction framework. We demonstrate the potential of FOF on both public datasets and real captured data. The code is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/FOF. Qiao Feng 0001, Yebin Liu, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
NeurIPS | 4 |
| 2022 | CdCLR: Clip-Driven Contrastive Learning for Skeleton-Based Action RecognitionabstractIn this study, we propose a Clip-Driven Contrastive Learning for Skeleton-Based Action Recognition (CdCLR). In-stead of considering sequences as instances, CdCLR extracts clips from the sequences as new instances. Aim to implement inherent supervision-guided contrastive learning through joint optimal training of sequences discrimination, clips discrimination, and order verification. Mining abundant positive/negative pairs inside sequence while learning inter-and intra-sequence semantic repre-sentations. Extensive experiments on the NTU RGB+D 60, UCLA and iMiGUE datasets present that CdCLR exhibits superior performance under various evaluation protocols and reaches state-of-the-art. Our code is available at https://github.com/Erich-G/CdCLRI. Rong Gao 0005, Xin Liu 0012, Jing-Yu Yang 0002, Huanjing Yue |
VCIP | 3 |
| 2022 | FPCR-Net: Feature pyramidal correlation and residual reconstruction for optical flow estimation
Jing-Yu Yang 0002, Cuiling Lan, Wenjun Zeng 0001 |
Neurocomputing | 3 |
| 2022 | Cloud Detection From Remote Sensing Imagery Based on Domain Translation NetworkabstractCloud detection in optical imagery has drawn remarkable attention in the era of big Earth observation data analytic. While multiple supervised learning models have been developed for such purpose, large volumes of paired training samples annotated at the pixel level are essential to ensure the model’s generalization capacity. However, constructing a comprehensive cloud detection training database is a tedious and time-consuming process. To tackle this dilemma, we simply regard cloud-contaminated remote sensing (RS) imagery as the combination of cloud and background domains and propose a cloud detection framework based on image-to-image domain translation network (DTNet) to separate cloud-contaminated RS imagery into two target domains of cloud and background object images without using any paired and pixel-level annotation training data. The framework was evaluated with multispectral images from two types of sensors, Landsat-8 Operational Land Imager (OLI) (30 m) and GaoFen-1 (16 m), and demonstrated superior or comparable performance compared with several state-of-the-art cloud detection models. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Yang Chen 0015, Chunping Hou, Kun Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Unsupervised Domain Adaptation for Cloud Detection Based on Grouped Features Alignment and Entropy MinimizationabstractMost convolutional neural network (CNN)-based cloud detection methods are built upon the supervised learning framework that requires a large number of pixel-level labels. However, it is expensive and time-consuming to manually annotate pixelwise labels for massive remote sensing images. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to generalize the model trained on labeled images of source satellite to unlabeled images of the target satellite. To effectively address the domain shift problem on cross-satellite images, we develop a novel UDA method based on grouped features alignment (GFA) and entropy minimization (EM) to extract domain-invariant representations to improve the cloud detection accuracy of cross-satellite images. The proposed UDA method is evaluated on “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks. Experimental results demonstrate the effectiveness of our method against existing state-of-the-art UDA approaches. The code of this paper has been made available online (https://github.com/nkszjx/grouped-features-alignment). Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Unsupervised Domain-Invariant Feature Learning for Cloud Detection of Remote Sensing ImagesabstractThe detection of clouds in remote sensing (RS) images is an important task, and convolutional neural networks (CNNs) have been used to perform it. However, supervised cloud detection CNNs rely heavily on a large number of samples annotated at pixel level to tune their parameter. Annotating RS images is a labor-intensive procedure and requires expert-level human knowledge. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to enable the model trained on labeled source satellite images to generalize to unlabeled target satellite images. Specifically, we propose a fine-grained feature alignment (FGFA) domain adaptation strategy that encourages a cloud detection network to extract domain-invariant representations, which improves the accuracy of cloud detection in unlabeled target satellite images. The proposed FGFA strategy consists of two steps: 1) fine-grained class-relevant feature selection based on an attention-guided mechanism and 2) class-relevant feature alignment (FA) based on a proposed grouped FA approach. Experimental results on the “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks demonstrate the effectiveness of our method and its superiority to existing state-of-the-art UDA approaches. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Xin Liu 0012, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Geometry-Guided Dense Perspective Network for Speech-Driven Facial AnimationabstractRealistic speech-driven 3D facial animation is a challenging problem due to the complex relationship between speech and face. In this paper, we propose a deep architecture, called Geometry-guided Dense Perspective Network (GDPnet), to achieve speaker-independent realistic 3D facial animation. The encoder is designed with dense connections to strengthen feature propagation and encourage the re-use of audio features, and the decoder is integrated with an attention mechanism to adaptively recalibrate point-wise feature responses by explicitly modeling interdependencies between different neuron units. We also introduce a non-linear face reconstruction representation as a guidance of latent space to obtain more accurate deformation, which helps solve the geometry-related deformation and is good for generalization across subjects. Huber and HSIC (Hilbert-Schmidt Independence Criterion) constraints are adopted to promote the robustness of our model and to better exploit the non-linear and high-order correlations. Experimental results on the public dataset and real scanned dataset validate the superiority of our proposed GDPnet compared with state-of-the-art model. The code is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/GDPnet. Jingying Liu, Binyuan Hui, Kun Li 0001, Yunke Liu, Yukun Lai, Yuxiang Zhang 0006, Yebin Liu, Jing-Yu Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2021 | Spatio-temporal Contrastive Domain Adaptation for Action RecognitionabstractCompared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling and alignment with frame-level or clip-level features, which is not discriminative sufficiently for video-based UDA tasks. To address these problems, in this paper we propose to establish the cross-modal domain alignment via self-supervised contrastive framework, i.e., spatio-temporal contrastive domain adaptation (STCDA), to learn the joint clip-level and video-level representation alignment. Since the effective representation is modeled from unlabeled data by self-supervised learning (SSL), spatio-temporal contrastive learning (STCL) is proposed to explore the useful long-term feature representation for classification, using self-supervision setting trained from the contrastive clip/video pairs with positive or negative properties. Besides, we involve a novel domain metric scheme, i.e., video-based contrastive alignment (VCA), to optimize the category-aware video-level alignment and generalization between source and target. The proposed STCDA achieves stat-of-the-art results on several UDA benchmarks for action recognition. Sicheng Zhao, Jing-Yu Yang 0002, Huanjing Yue, Pengfei Xu 0013, Runbo Hu |
CVPR | 3 |
| 2021 | PISE: Person Image Synthesis and Editing With Decoupled GANabstractPerson image synthesis, e.g., pose transfer, is a challenging problem due to large variation and occlusion. Existing methods have difficulties predicting reasonable invisible regions and fail to decouple the shape and style of clothing, which limits their applications on person image editing. In this paper, we propose PISE, a novel two-stage generative model for Person Image Synthesis and Editing, which is able to generate realistic person images with desired poses, textures, or semantic layouts. For human pose transfer, we first synthesize a human parsing map aligned with the target pose to represent the shape of clothing by a parsing generator, and then generate the final image by an image generator. To decouple the shape and style of clothing, we propose joint global and local per-region encoding and normalization to predict the reasonable style of clothing for invisible regions. We also propose spatial-aware normalization to retain the spatial context relationship in the source image. The results of qualitative and quantitative experiments demonstrate the superiority of our model on human pose transfer. Besides, the results of texture transfer and region editing show that our model can be applied to person image editing. The code is available for research purposes at https://github.com/Zhangjinso/PISE. Kun Li 0001, Yukun Lai, Jing-Yu Yang 0002 |
CVPR | 4 |
| 2021 | Implicit Transformer Network for Screen Content Image Continuous Super-ResolutionabstractNowadays, there is an explosive growth of screen contents due to the wide application of screen sharing, remote cooperation, and online education. To match the limited terminal bandwidth, high-resolution (HR) screen contents may be downsampled and compressed. At the receiver side, the super-resolution (SR)of low-resolution (LR) screen content images (SCIs) is highly demanded by the HR display or by the users to zoom in for detail observation. However, image SR methods mostly designed for natural images do not generalize well for SCIs due to the very different image characteristics as well as the requirement of SCI browsing at arbitrary scales. To this end, we propose a novel Implicit Transformer Super-Resolution Network (ITSRN) for SCISR. For high-quality continuous SR at arbitrary ratios, pixel values at query coordinates are inferred from image features at key coordinates by the proposed implicit transformer and an implicit position encoding scheme is proposed to aggregate similar neighboring pixel values to the query one. We construct benchmark SCI1K and SCI1K-compression datasets withLR and HR SCI pairs. Extensive experiments show that the proposed ITSRN significantly outperforms several competitive continuous and discrete SR methods for both compressed and uncompressed SCIs. Jing-Yu Yang 0002, Sheng Shen 0010, Huanjing Yue, Kun Li 0001 |
NeurIPS | 1 |
| 2021 | Low-light image enhancement based on Retinex decomposition and adaptive gamma correctionabstractAbstract Low‐light images suffer from poor visibility and noise. In this paper, a low‐light image enhancement method based on Retinex decomposition is proposed. A pyramid network is first utilized to extract multi‐scale features to improve the quality of Retinex decomposition. Then the decomposed illumination is refined via an adaptive Gamma correction network to handle non‐uniform illumination, while the decomposed reflectance is refined with a lightweight network. Finally, the enhanced image is obtained by element‐wise multiplication between the refined illumination and reflectance components. Quantitative and qualitative experiments demonstrate the superiority of our method over state‐of‐the‐art image enhancement methods. Jing-Yu Yang 0002, Huanjing Yue, Zhongyu Jiang, Kun Li 0001 |
IET Image Process. | 1 |
| 2021 | Unsupervised moiré pattern removal for recaptured screen images
Huanjing Yue, Yijia Cheng, Fanglong Liu, Jing-Yu Yang 0002 |
Neurocomputing | 4 |
| 2021 | Deep edge map guided depth super resolution
Zhongyu Jiang, Huanjing Yue, Yukun Lai, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou |
Signal Process. Image Commun. | 4 |
| 2021 | Sparse intrinsic decomposition and applications
Kun Li 0001, Xinchen Ye, Chenggang Yan 0001, Jing-Yu Yang 0002 |
Signal Process. Image Commun. | 5 |
| 2021 | STC-Flow: Spatio-temporal context-aware optical flow estimation
Jing-Yu Yang 0002 |
Signal Process. Image Commun. | 3 |
| 2021 | Deep noise estimation and removal for real-world noisy images
Huanjing Yue, Zhongyu Jiang, Shengdi Zhou, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou |
Signal Process. Image Commun. | 4 |
| 2021 | Reference guided image super-resolution via efficient dense warping and adaptive fusion
Huanjing Yue, Zhongyu Jiang, Jing-Yu Yang 0002, Chunping Hou |
Signal Process. Image Commun. | 4 |
| 2021 | Recaptured Screen Image DemoiréingabstractIn many situations, such as transferring data between devices and recording precious moments, we would like to capture the contents on screens using digital cameras for convenience. These recaptured screen images and videos suffer from a special type of degradation called “moiré pattern”, which is caused by the aliasing between the grid of display screen and the array of camera sensor. However, few works are proposed to tackle this problem. Considering the great success of convolutional neural networks (CNNs) in image restoration, we propose a CNN-based moiré removal method for recaptured screen images. There are mainly two contributions in this paper. First, for the generation of training data, we propose an image registration algorithm via global homography transform and local patch matching to compensate the significant viewpoint disparity between the recaptured screen image and the moiré-free image obtained via screenshot. We construct a moiré removal and brightness improvement (MRBI) database with aligned moiré-free and moiré images. Second, we propose a convolutional neural Network with Additive and Multiplicative modules (termed as AMNet) to transfer the low light moiré image to the bright moiré-free image. The proposed network is trained with pixel-wise loss, perceptual loss, and adversarial loss. Extensive experiments on 340 test images demonstrate that the proposed method outperforms state-of-the-art moiré removal methods. Huanjing Yue, Lipu Liang, Hongteng Xu, Chunping Hou, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Landsat-8 OLI Multispectral Image Dehazing Based on Optimized Atmospheric Scattering ModelabstractOptical satellite images are often affected by haze atmospheric conditions, which degrades the quality of remote sensing (RS) data and reduces the accuracy of interpretation and classification. Hence, haze removal becomes a necessary preprocessing step for most of the applications of RS image. In this article, we propose a novel haze removal method for Landsat-8 OLI multispectral image based on an optimized atmospheric scattering model. We focus on adaptively estimating the haze transmission map of each band by taking into account the effect of both wavelength and haze atmospheric conditions (haze particle size and haze particle concentration) thus improving dehazing performance. The experimental results on Landsat-8 OLI multispectral images show that the proposed dehazing model is able to remove haze successfully and significantly improve the image visibility as well as correct the spectral bias to some degree. Moreover, this method is simple and feasible, and has good practical value. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | CDnetV2: CNN-Based Cloud Detection for Remote Sensing Imagery With Cloud-Snow CoexistenceabstractCloud detection is a crucial preprocessing step for optical satellite remote sensing (RS) images. This article focuses on the cloud detection for RS imagery with cloud-snow coexistence and the utilization of the satellite thumbnails that lose considerable amount of high resolution and spectrum information of original RS images to extract cloud mask efficiently. To tackle this problem, we propose a novel cloud detection neural network with an encoder-decoder structure, named CDnetV2, as a series work on cloud detection. Compared with our previous CDnetV1, CDnetV2 contains two novel modules, that is, adaptive feature fusing model (AFFM) and high-level semantic information guidance flows (HSIGFs). AFFM is used to fuse multilevel feature maps by three submodules: channel attention fusion model (CAFM), spatial attention fusion model (SAFM), and channel attention refinement model (CARM). HSIGFs are designed to make feature layers at decoder of CDnetV2 be aware of the locations of the cloud objects. The high-level semantic information of HSIGFs is extracted by a proposed high-level feature fusing model (HFFM). By being equipped with these two proposed key modules, AFFM and HSIGFs, CDnetV2 is able to fully utilize features extracted from encoder layers and yield accurate cloud detection results. Experimental results on the ZY-3 satellite thumbnail data set demonstrate that the proposed CDnetV2 achieves accurate detection accuracy and outperforms several state-of-the-art methods. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | RSDehazeNet: Dehazing Network With Channel Refinement for Multispectral Remote Sensing ImagesabstractMultispectral remote sensing (RS) images are often contaminated by the haze that degrades the quality of RS data and reduces the accuracy of interpretation and classification. Recently, the emerging deep convolutional neural networks (CNNs) provide us new approaches for RS image dehazing. Unfortunately, the power of CNNs is limited by the lack of sufficient hazy-clean pairs of RS imagery, which makes supervised learning impractical. To meet the data hunger of supervised CNNs, we propose a novel haze synthesis method to generate realistic hazy multispectral images by modeling the wavelength-dependent and spatial-varying characteristics of haze in RS images. The proposed haze synthesis method not only alleviates the lack of realistic training pairs in multispectral RS image dehazing but also provides a benchmark data set for quantitative evaluation. Furthermore, we propose an end-to-end RSDehazeNet for haze removal. We utilize both local and global residual learning strategies in RSDehazeNet for fast convergence with superior performance. Channel attention modules are incorporated to exploit strong channel correlation in multispectral RS images. Experimental results show that the proposed network outperforms the state-of-the-art methods for synthetic data and real Landsat-8 OLI multispectral RS images. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Supervised Raw Video Denoising With a Benchmark Dataset on Dynamic ScenesabstractIn recent years, the supervised learning strategy for real noisy image denoising has been emerging and has achieved promising results. In contrast, realistic noise removal for raw noisy videos is rarely studied due to the lack of noisy-clean pairs for dynamic scenes. Clean video frames for dynamic scenes cannot be captured with a long-exposure shutter or averaging multi-shots as was done for static images. In this paper, we solve this problem by creating motions for controllable objects, such as toys, and capturing each static moment for multiple times to generate clean video frames. In this way, we construct a dataset with 55 groups of noisy-clean videos with ISO values ranging from 1600 to 25600. To our knowledge, this is the first dynamic video dataset with noisy-clean pairs. Correspondingly, we propose a raw video denoising network (RViDeNet) by exploring the temporal, spatial, and channel correlations of video frames. Since the raw video has Bayer patterns, we pack it into four sub-sequences, i.e RGBG sequences, which are denoised by the proposed RViDeNet separately and finally fused into a clean video. In addition, our network not only outputs a raw denoising result, but also the sRGB result by going through an image signal processing (ISP) module, which enables users to generate the sRGB result with their favourite ISPs. Experimental results demonstrate that our method outperforms state-of-the-art video and raw image denoising algorithms on both indoor and outdoor videos. Huanjing Yue, Cong Cao 0005, Ronghe Chu, Jing-Yu Yang 0002 |
CVPR | 5 |
| 2020 | 3D Motion Recovery via Low Rank Matrix Restoration with Hankel-Like AugmentationabstractThis paper proposes a 3D skeleton recovery model equipped with a joint augmented low-rank and sparse prior and an articulation-graph-based isometric constraint to exploit temporal and spatial correlation, respectively. The corrupted 3D skeleton sequence is represented as a matrix and constrained by several priors in the proposed model. A Hankel-like augmentation is adopted to strengthen the low-rankness and we integrate a decoupling technique to reduce the internal interferences of the data. We solve our model via an alternating direction method under the augmented Lagrangian multiplier framework with a Gauss-Newton solver for the subproblem of isometric optimization. Experimental results on two skeleton datasets demonstrate the effectiveness and superiority of the proposed model in motion reconstruction and skeleton recovery, compared with state-of-the-art methods. Jing-Yu Yang 0002, Jiabin Shi, Yuyuan Zhu, Kun Li 0001, Chunping Hou |
ICME | 1 |
| 2020 | Depth Super-Resolution via Deep Controllable Slicing NetworkabstractDue to the imaging limitation of depth sensors, high-resolution (HR) depth maps are often difficult to be acquired directly, thus effective depth super-resolution (DSR) algorithms are needed to generate HR output from its low-resolution (LR) counterpart. Previous methods treat all depth regions equally without considering different extents of degradation at region-level, and regard DSR under different scales as independent tasks without considering the modeling of different scales, which impede further performance improvement and practical use of DSR. To alleviate these problems, we propose a deep controllable slicing network from a novel perspective. Specifically, our model is to learn a set of slicing branches in a divide-and-conquer manner, parameterized by a distance-aware weighting scheme to adaptively aggregate different depths in an ensemble. Each branch that specifies a depth slice (e.g., the region in some depth range) tends to yield accurate depth recovery. Meanwhile, a scale-controllable module that extracts depth features under different scales is proposed and inserted into the front of slicing network, and enables finely-grained control of the depth restoration results of slicing network with a scale hyper-parameter. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our method achieves superior performance. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
ACM Multimedia | 4 |
| 2020 | SHREC'20: Shape correspondence with non-isometric deformations abstractEstimating correspondence between two shapes continues to be a challenging problem in geometry processing. Most current methods assume deformation to be near-isometric, however this is often not the case. For this paper, a collection of shapes of different animals has been curated, where parts of the animals (e.g., mouths, tails & ears) correspond yet are naturally non-isometric. Ground-truth correspondences were established by asking three specialists to independently label corresponding points on each of the models with respect to a previously labelled reference model. We employ an algorithmic strategy to select a single point for each correspondence that is representative of the proposed labels. A novel technique that characterises the sparsity and distribution of correspondences is employed to measure the performance of ten shape correspondence methods. Roberto M. Dyke, Yukun Lai, Paul L. Rosin, Stefano Zappalà, Seana Dykes, Daoliang Guo, Kun Li 0001, Riccardo Marin, Simone Melzi, Jing-Yu Yang 0002 |
Comput. Graph. | 10 |
| 2020 | Reference Image Guided Super-Resolution via Progressive Channel Attention Networks
Huanjing Yue, Sheng Shen 0010, Jing-Yu Yang 0002, Haofeng Hu, Yan-Fang Chen |
J. Comput. Sci. Technol. | 3 |
| 2020 | Depth upsampling based on deep edge-aware learning
Zhihui Wang 0001, Xinchen Ye, Baoli Sun, Jing-Yu Yang 0002, Rui Xu 0002 |
Pattern Recognit. | 4 |
| 2020 | A sparsity-promoting image decomposition model for depth recovery
Xinchen Ye, Mingliang Zhang 0002, Jing-Yu Yang 0002, Xin Fan 0001, Fangfang Guo |
Pattern Recognit. | 3 |
| 2020 | Adversarial unsupervised domain adaptation for cross scenario waveform recognition
Qing Wang 0015, Panfei Du, Xiaofeng Liu 0009, Jing-Yu Yang 0002, Guohua Wang 0002 |
Signal Process. | 4 |
| 2020 | Spatiotemporally scalable matrix recovery for background modeling and moving object detection
Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou |
Signal Process. | 1 |
| 2020 | Temporal-Spatial Mapping for Action RecognitionabstractDeep learning models have enjoyed great success for image related computer vision tasks such as image classification and object detection. For video related tasks such as human action recognition, however, the advancements are not as significant yet. The main challenge is the lack of effective and efficient models in modeling the rich temporal-spatial information in a video. We introduce a simple yet effective operation, termed temporal-spatial mapping, for capturing the temporal evolution of the frames by jointly analyzing all the frames of a video. We propose a video level 2D feature representation by transforming the convolutional features of all frames to a 2D feature map, referred to as VideoMap. With each row being the vectorized feature representation of a frame, the temporal-spatial features are compactly represented, while the temporal dynamic evolution is also well embedded. Based on the VideoMap representation, we further propose a temporal attention model within a shallow convolutional neural network to efficiently exploit the temporal-spatial dynamics. The experiment results show that the proposed scheme achieves state-of-the-art performance, with 4.2% accuracy gain over the temporal segment network, a competing baseline method, on the challenging human action benchmark dataset HMDB51. Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Xiaoyan Sun 0001, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Spatio-Temporal Reconstruction for 3D Motion RecoveryabstractThis paper addresses the challenge of 3D motion recovery by exploiting the spatio-temporal correlations of corrupted 3D skeleton sequences. We propose a new 3D motion recovery method using spatio-temporal reconstruction, which uses joint low-rank and sparse priors to exploit temporal correlation and an isometric constraint for spatial correlation. The proposed model is formulated as a constrained optimization problem, which is efficiently solved by the augmented Lagrangian method with a Gauss-Newton solver for the subproblem of isometric optimization. The experimental results on the CMU motion capture dataset, Edinburgh dataset, and two Kinect datasets demonstrate that the proposed approach achieves better motion recovery than the state-of-the-art methods. The proposed method is applicable to Kinect-like skeleton tracking devices and pose estimation methods that cannot provide accurate estimation of complex motions, especially in the presence of occlusion. Jing-Yu Yang 0002, Kun Li 0001, Meiyuan Wang, Yukun Lai, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | IENet: Internal and External Patch Matching ConvNet for Web Image Guided DenoisingabstractFrom the non-local self-similarity (NSS)-based image denoising to the convolutional-network (ConvNet)-based image denoising, the denoising performance has been greatly improved. However, it is still not clear how to utilize similar web images to guide image denoising using ConvNet. This paper proposes a novel ConvNet for image denoising to explore both internal (NSS) and external correlations when external similar images are available. Since external similar images may be taken with different viewpoints, focal lengths, and may contain different objects, it is difficult to directly explore external correlations at image level using ConvNet. Therefore, we propose an internal and external patch matching ConvNet (IENet), whose inputs are similar patch cubes extracted from the noisy input and its external similar images. We design three different network structures, namely early-fusion, middle-fusion, and late-fusion of the internal and external cubes to fully combine the strengths of internal and external correlations. The experimental results demonstrate that the proposed method achieves the best denoising results compared with the seven state-of-the-art denoising methods. In specific, the proposed method outperforms the state-of-the-art web image guided denoising method by more than 1 dB on average, which further demonstrates the superiority of the proposed IENet-based filtering over the hand-crafted filtering methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Truong Q. Nguyen, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Discern Depth Under Foul Weather: Estimate PM2.5 for Depth InferenceabstractNowadays, haze is a common and serious problem and PM$_{2.5}$is a main measurement for air quality. Current methods estimate the level of primary pollutant with professional instruments, which is expensive and inconvenient. Moreover, with haze, the captured images will be unclear and are difficult to estimate the depth of the scene using passive methods. This article proposes a cheap, fast, and convenient PM$_{2.5}$estimation method that only need a captured image using daily-life devices, and further, discerns the depth of the scene using the estimated PM$_{2.5}$. We learn haze-relevant classified mapping via the hybrid convolutional neural network and combine the high-level features extracted from the convolutional layer with ground-truth PM$_{2.5}$to train support vector regression. The transmission map is computed using nonlocal sparse priors, and the depth map is inferred using the estimated PM$_{2.5}$value through the atmospheric scattering model. Experimental results demonstrate that our method achieves accurate PM$_{2.5}$estimation and depth inference. This could be very useful in many applications, for both clean and foul weather. Kun Li 0001, Yahong Han, Xibin Yue, Jing-Yu Yang 0002 |
IEEE Trans. Ind. Informatics | 7 |
| 2020 | Learning to Reconstruct and Understand Indoor Scenes From Sparse ViewsabstractThis paper proposes a new method for simultaneous 3D reconstruction and semantic segmentation for indoor scenes. Unlike existing methods that require recording a video using a color camera and/or a depth camera, our method only needs a small number of (e.g., 3~5) color images from uncalibrated sparse views, which significantly simplifies data acquisition and broadens applicable scenarios. To achieve promising 3D reconstruction from sparse views with limited overlap, our method first recovers the depth map and semantic information for each view, and then fuses the depth maps into a 3D scene. To this end, we design an iterative deep architecture, named IterNet, to estimate the depth map and semantic segmentation alternately. To obtain accurate alignment between views with limited overlap, we further propose a joint global and local registration method to reconstruct a 3D scene with semantic information. We also make available a new indoor synthetic dataset, containing photorealistic high-resolution RGB images, accurate depth maps and pixel-level semantic labels for thousands of complex layouts. Experimental results on public datasets and our dataset demonstrate that our method achieves more accurate depth estimation, smaller semantic segmentation errors, and better 3D reconstruction results over state-of-the-art methods. Jing-Yu Yang 0002, Kun Li 0001, Yukun Lai, Huanjing Yue, Jianzhi Lu, Hao Wu 0042, Yebin Liu |
IEEE Trans. Image Process. | 1 |
| 2020 | PMBANet: Progressive Multi-Branch Aggregation Network for Scene Depth Super-ResolutionabstractDepth map super-resolution is an ill-posed inverse problem with many challenges. First, depth boundaries are generally hard to reconstruct particularly at large magnification factors. Second, depth regions on fine structures and tiny objects in the scene are destroyed seriously by downsampling degradation. To tackle these difficulties, we propose a progressive multi-branch aggregation network (PMBANet), which consists of stacked MBA blocks to fully address the above problems and progressively recover the degraded depth map. Specifically, each MBA block has multiple parallel branches: 1) The reconstruction branch is proposed based on the designed attention-based error feed-forward/-back modules, which iteratively exploits and compensates the downsampling errors to refine the depth map by imposing the attention mechanism on the module to gradually highlight the informative features at depth boundaries. 2) We formulate a separate guidance branch as prior knowledge to help to recover the depth details, in which the multi-scale branch is to learn a multi-scale representation that pays close attention at objects of different scales, while the color branch regularizes the depth map by using auxiliary color information. Then, a fusion block is introduced to adaptively fuse and select the discriminative features from all the branches. The design methodology of our whole network is well-founded, and extensive experiments on benchmark datasets demonstrate that our method achieves superior performance in comparison with the state-of-the-art methods. Our code and models are available athttps://github.com/Sunbaoli/PMBANet_DSR/. Xinchen Ye, Baoli Sun, Zhihui Wang 0001, Jing-Yu Yang 0002, Rui Xu 0002, Baopu Li |
IEEE Trans. Image Process. | 4 |
| 2019 | Graph Based Non-Uniform Sampling and Reconstruction of Depth MapsabstractHigh-quality depth sensing is highly demanded in intelligent computer vision, 3DTV, and many other related fields. However, prevalent time-of-fly (ToF) depth sensors are of low resolution as the number of pixel-level demodulators is limited. Moreover, the rectangular sampling does not consider the signal characteristics of depth maps. Being a departure of previous resolution enhancement on rectangular sampling, this paper investigates the non-uniform sampling of depth maps, and the high-resolution depth reconstruction from limited non-uniformly distributed samples. The proposed depth sampling and reconstruction schemes are developed based on graph signal processing. We first propose a graph-based non-uniform sampling (GNS) scheme, where depth signals are sampled based on the response of a high-pass graph filter, which results in denser sampling around discontinuities such as edges and contours than in smooth regions. We then propose a graph-based depth reconstruction (GDR) framework,where a graph Laplacian regularizer is designed to fully exploit structural correlation between the depth and photometric images. To solve the reconstruction problem, we derive an efficient algorithm based on the alternating direction method of multipliers (ADMM). Experimental results show that the GNS-GDR non-uniform sampling and reconstruction method achieves high-quality depth sensing, outperforming several state-of-the-art schemes. Jing-Yu Yang 0002, Xinchen Ye, Pascal Frossard, Kun Li 0001 |
ICIP | 1 |
| 2019 | Global as-Conformal-as-Possible Non-Rigid Registration of Multi-view ScansabstractIn this paper, we present a novel framework for global non-rigid registration of multi-view scans captured using consumer-level depth cameras. In our method, all scans from different viewpoints are allowed to undergo large non-rigid deformations and finally fused into a complete high quality model. To avoid the well-known loop closure problem, we simultaneously optimize a global alignment problem instead of pairwise non-rigid registration in succession. We employ a joint point-to-point and point-to-plane positional constraint to reduce the influence of wrong correspondences, and incorporate an as-conformal-as-possible constraint to avoid mesh distortions during deformation. We also design a reweighting scheme on position and transformation to reduce registration errors. Experimental results on both public datasets and real scanned datasets demonstrate that our approach outperforms state-of-the-art methods through extensive quantitative and qualitative evaluations. Zhenchao Wu, Kun Li 0001, Yukun Lai, Jing-Yu Yang 0002 |
ICME | 4 |
| 2019 | 3D Mesh Based Inter-Image Prediction for Image Set CompressionabstractA key problem in image set compression is inter-image prediction. Different from the conventional 2D transformation based methods, in this paper we propose a novel 3D mesh based inter-image prediction method. We reconstruct a 3D mesh from the images in the set as a compact representation of the photographed scene. Regarding the images as different projections of the mesh, we build coordinates mappings between images by the multi-view geometry. Exploiting the continuity of the mesh surface, we naturally model the occlusions in the scene and perform inter-image prediction with higher accuracy. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts significantly. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
ICME | 3 |
| 2019 | 3D Face Reprentation and Reconstruction with Multi-scale Graph Convolutional AutoencodersabstractEffective representation and reconstruction for human faces are very important in many applications. Existing linear representation methods cannot reconstruct high quality 3D faces with details, while the newest non-linear representation method is less suitable for real shapes since spectral decompositions are unstable across different graphs. To address these problems, we propose a multi-scale graph convolutional autoencoder for face representation and reconstruction. Our autoencoder uses graph convolution, which is easily trained for the data with graph structures and can be used for other deformable models. Our model can also be used for variational training to generate high quality face shapes. Experimental results demonstrate that our model can generate more plausible, complex, and stable 3D shapes, and achieves higher quality face reconstruction compared with state-of-the-art methods. Cunkuan Yuan, Kun Li 0001, Yukun Lai, Yebin Liu, Jing-Yu Yang 0002 |
ICME | 5 |
| 2019 | Generating 3D Faces using Multi-column Graph Convolutional NetworksabstractAbstract In this work, we introduce multi‐column graph convolutional networks (MGCNs), a deep generative model for 3D mesh surfaces that effectively learns a non‐linear facial representation. We perform spectral decomposition of meshes and apply convolutions directly in the frequency domain. Our network architecture involves multiple columns of graph convolutional networks (GCNs), namely large GCN (L‐GCN), medium GCN (M‐GCN) and small GCN (S‐GCN), with different filter sizes to extract features at different scales. L‐GCN is more useful to extract large‐scale features, whereas S‐GCN is effective for extracting subtle and fine‐grained features, and M‐GCN captures information in between. Therefore, to obtain a high‐quality representation, we propose a selective fusion method that adaptively integrates these three kinds of information. Spatially non‐local relationships are also exploited through a self‐attention mechanism to further improve the representation ability in the latent vector space. Through extensive experiments, we demonstrate the superiority of our end‐to‐end framework in improving the accuracy of 3D face reconstruction. Moreover, with the help of variational inference, our model has excellent generating ability. Kun Li 0001, Jingying Liu, Yukun Lai, Jing-Yu Yang 0002 |
Comput. Graph. Forum | 4 |
| 2019 | Transferred deep learning based waveform recognition for cognitive passive radar
Qing Wang 0015, Panfei Du, Jing-Yu Yang 0002, Guohua Wang 0002, Jianjun Lei 0001, Chunping Hou |
Signal Process. | 3 |
| 2019 | CDnet: CNN-Based Cloud Detection for Remote Sensing ImageryabstractCloud detection is one of the important tasks for remote sensing image (RSI) preprocessing. In this paper, we utilize the thumbnail (i.e., preview image) of RSI, which contains the information of original multispectral or panchromatic imagery, to extract cloud mask efficiently. Compared with detection cloud mask from original RSI, it is more challenging to detect cloud mask using thumbnails due to the loss of resolution and spectrum information. To tackle this problem, we propose a cloud detection neural network (CDnet) with an encoder-decoder structure, a feature pyramid module (FPM), and a boundary refinement (BR) block. The FPM extracts the multiscale contextual information without the loss of resolution and coverage; the BR block refines object boundaries; and the encoder-decoder structure gradually recovers segmentation results with the same size as input image. Experimental results on the ZY-3 satellite thumbnails cloud cover validation data set and two other validation data sets (GF-1 WFV Cloud and Cloud Shadow Cover Validation Data and Landsat-8 Cloud Cover Assessment Validation Data) demonstrate that the proposed method achieves accurate detection accuracy and outperforms several state-of-the-art methods. Jing-Yu Yang 0002, Jianhua Guo 0002, Huanjing Yue, Haofeng Hu, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Global 3D Non-Rigid Registration of Deformable Objects Using a Single RGB-D CameraabstractWe present a novel global non-rigid registration method for dynamic 3D objects. Our method allows objects to undergo large non-rigid deformations and achieves high-quality results even with substantial pose change or camera motion between views. In addition, our method does not require a template prior and uses less raw data than tracking-based methods since only a sparse set of scans is needed. We simultaneously compute the deformations of all the scans by optimizing a global alignment problem to avoid the well-known loop closure problem and use an as-rigid-as-possible constraint to eliminate the shrinkage problem of the deformed shapes, especially near open boundaries of scans. To cope with large-scale problems, we design a coarse-to-fine multi-resolution scheme, which also avoids the optimization being trapped into local minima. The proposed method is evaluated on public datasets and real datasets captured by an RGB-D sensor. The experimental results demonstrate that the proposed method obtains better results than several state-of-the-art methods. Jing-Yu Yang 0002, Daoliang Guo, Kun Li 0001, Zhenchao Wu, Yukun Lai |
IEEE Trans. Image Process. | 1 |
| 2019 | High ISO JPEG Image Denoising by Deep Fusion of Collaborative and Convolutional FilteringabstractCapturing images at high ISO modes will introduce much realistic noise, which is difficult to be removed by traditional denoising methods. In this paper, we propose a novel denoising method for high ISO JPEG images via deep fusion of collaborative and convolutional filtering. Collaborative filtering explores the non-local similarity of natural images, while convolutional filtering takes advantage of the large capacity of convolutional neural networks (CNNs) to infer noise from noisy images. We observe that the noise variance map of a high ISO JPEG image is spatial-dependent and has a Bayer-like pattern. Therefore, we introduce the Bayer pattern prior in our noise estimation and collaborative filtering stages. Since collaborative filtering is good at recovering repeatable structures and convolutional filtering is good at recovering irregular patterns and removing noise in flat regions, we propose to fuse the strengths of the two methods via deep CNN. The experimental results demonstrate that our method outperforms the state-of-the-art realistic noise removal methods for a wide variety of testing images in both subjective and objective measurements. In addition, we construct a dataset with noisy and clean image pairs for high ISO JPEG images to facilitate research on this topic. Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Robust Non-Rigid Registration with Reweighted Position and Transformation SparsityabstractNon-rigid registration is challenging because it is ill-posed with high degrees of freedom and is thus sensitive to noise and outliers. We propose a robust non-rigid registration method using reweighted sparsities on position and transformation to estimate the deformations between 3-D shapes. We formulate the energy function with position and transformation sparsity on both the data term and the smoothness term, and define the smoothness constraint using local rigidity. The double sparsity based non-rigid registration model is enhanced with a reweighting scheme, and solved by transferring the model into four alternately-optimized subproblems which have exact solutions and guaranteed convergence. Experimental results on both public datasets and real scanned datasets show that our method outperforms the state-of-the-art methods and is more robust to noise and outliers than conventional non-rigid registration methods. Kun Li 0001, Jing-Yu Yang 0002, Yukun Lai, Daoliang Guo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography EstimationabstractIt is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT) features of the images, and then adopt the multi-homography fitting algorithm to classify the feature points into different deformation models. According to the deduced deformation models, we partition the source image into base and transition regions. For the base regions, we adopt the moving direct linear transformation (Moving DLT) to estimate homographies. For the transition regions, we propose a hierarchical homography estimation method to select appropriate homographies. Experimental results show that our method achieves more accurate alignment results compared with state-of-the-art alignment methods for building images. Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou |
ICASSP | 2 |
| 2018 | Image-Based PM2.5 Estimation and its Application on Depth EstimationabstractAir pollution is still a big threat to human health particularly for developing countries. It is highly demanding to measure air quality with daily-used devices such as smartphones. On the other hand, it is difficult to estimate the scene depth under the foul weather using traditional vision-based methods. This paper proposes an image-based method for PM2.5 estimation by capturing a single image. We extract high-level features based on convolutional neural network (CNN) and learn the mapping between the features and PM2.5 by support vector regression (SVR). Given a captured image, we can estimate the PM2.5 value in real time. With the estimated PM2.5, we can estimate the depth of scene using sparse prior and nonlocal bilateral kernel. Experimental results demonstrate that the proposed method achieves the same accuracy of PM2.5 estimation as commodity measurement devices, and estimates the accurate depth information that is even better than the “ground-truth” captured by a laser in the no-haze condition. Kun Li 0001, Yahong Han, Pufeng Du, Jing-Yu Yang 0002 |
ICASSP | 5 |
| 2018 | Image-based Air Pollution Estimation Using Hybrid Convolutional Neural NetworkabstractAir pollution has a serious impact on our daily life, and how to quickly and easily measure the air pollution level without any expensive equipment is a quite challenging task. This paper proposes an air pollution estimation method using deep hybrid convolutional neural network from a single image, e.g., captured by a smartphone. The captured image is input to the main network, a very deep network, which solves the side effects of increased depth (degradation issues) by skip connection. This can improve network performance by simply increasing the depth of the network. Dark channel map is computed and fed into a secondary network to enrich the features with implicit representation. We have collected 1575 images of different scenes with different values of PM2.5to train the network in the end-to-end fusion mode. Experimental results on synthetic dataset and real captured dataset demonstrate that our method achieves excellent performance on classification of air pollution levels from a single captured image. Kun Li 0001, Yahong Han, Jing-Yu Yang 0002 |
ICPR | 4 |
| 2018 | Deep Joint Noise Estimation and Removal for High ISO JPEG ImagesabstractCapturing images under high ISO mode introduces much noise. The statistics of high ISO noise is quite different from that of Gaussian noise. Therefore, this kind of noise is difficult to be removed by traditional Gaussian noise removal methods. This paper proposes a convolutional neural network (CNN) based method to jointly estimate and remove high ISO noise. There are two contributions in this paper. First, we propose a CNN based noise estimation method to estimate the pixel-wise noise level. Due to the Bayer down-sampling process in imaging, the noise variance map is characterized by Bayer patterns. Therefore, we propose packing 2 × 2 blocks in a noisy image into 4D vectors, which makes the pixels with similar noise levels be neighbors. Second, the noise variance map is correlated with the image content. Thus, we propose concatenating the estimated noise variance map with the noisy image, and feed the fused data to the denoising network. The two networks are trained together in an end-to-end fashion. Experimental results demonstrate that the proposed method outperforms state-of-the-art noise estimation and removal methods. Huanjing Yue, Shengdi Zhou, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Chunping Hou |
ICPR | 3 |
| 2018 | Shape and Pose Estimation for Closely Interacting Persons Using Multi-view ImagesabstractAbstract Multi‐person pose and shape estimation is very challenging, especially when the persons have close interactions. Existing methods only work well when people are well spaced out in the captured images. However, close interaction among people is very common in real life, which is more challenge due to complex articulation, frequent occlusion and inherent ambiguities. We present a fully‐automatic markerless motion capture method to simultaneously estimate 3D poses and shapes of closely interacting people from multi‐view sequences. We first predict the 2D joints for each person in an image, and then design a spatio‐temporal tracker for multi‐person pose tracking based on multi‐view videos. Finally, we estimate 3D poses and shapes of all the persons with multi‐view constraints using a skinned multi‐person linear model (SMPL). Experimental results demonstrate that our method achieves fast but accurate pose and shape estimation results for multi‐person close interaction cases. Compared with existing methods, our method does not need pre‐segmentation for each person and manual intervention, which greatly reduces the complexity of the system including time complexity and system processing complexity. Kun Li 0001, Nianhong Jiao, Yebin Liu, Yangang Wang 0001, Jing-Yu Yang 0002 |
Comput. Graph. Forum | 5 |
| 2018 | Non-orthogonal frequency division multiplexing based on sparse representationabstractThis study proposes a sparse non‐orthogonal frequency division multiplexing (S‐NOFDM) based on sparse representation to improve the spectral efficiency of orthogonal frequency division multiplexing (OFDM). The subcarriers of S‐NOFDM are generated from a subset of OFDM orthogonal subcarriers, which therefore requires less spectral resource. Each selected OFDM subcarrier is shifted into a group of overlapping subcarriers with different time delays. The modulated signal is produced by solving the sparse representation of the input signal under the generated set of subcarriers. The demodulation is simply a linear combination of generated subcarriers with the recovered modulated signal. Simulation results show that the proposed S‐NOFDM can achieve better bit error rate performance in an additive white Gaussian noise channel and a little worse than OFDM in a Rayleigh channel, with less frequency resources required. Xiaomei Fu, Jing-Yu Yang 0002 |
IET Commun. | 3 |
| 2018 | Depth Super-Resolution From RGB-D Pairs With Transform and Spatial Domain RegularizationabstractThis paper proposes a depth super-resolution method with both transform and spatial domain regularization. In the transform domain regularization, nonlocal correlations are exploited via an auto-regressive model, where each patch is further sparsified with a locally-trained transform to consider intra-patch correlations. In the spatial domain regularization, we propose a multi-directional total variation (MTV) prior to characterize the geometrical structures spatially orientated at arbitrary directions in depth maps. To achieve adaptive regularization, the MTV is weighted for each directional finite difference considering local characteristics of RGB-D data. We develop an accelerated proximal gradient algorithm to solve the proposed model. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior depth super-resolution performance for various configurations of magnification factors and datasets. Zhongyu Jiang, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002, Chunping Hou |
IEEE Trans. Image Process. | 4 |
| 2017 | Estimation of signal-dependent noise level function using multi-column convolutional neural networkabstractTo estimate the levels of signal-dependent noise (SDN) from a single image is challenging. This paper proposes a novel method to estimate the noise level function (NLF) from a single image using a Multi-column Convolutional Neural Network (MC-Net) with an end-to-end architecture. The MC-Net is trained on a synthesized dataset containing noisy images with known NLFs, and it allows to learn rich hierarchical features using three sub-networks. Moreover, this method performs end-to-end training to retain more details for pixel-wise noise level estimation. Experimental results indicate that our method is accurate and robust to estimate NLFs of SDN for various types of images. Jing-Yu Yang 0002, Xin Liu 0012, Kun Li 0001 |
ICIP | 1 |
| 2017 | Underwater image enhancement based on structure-texture decompositionabstractUnderwater images generally suffer from low contrast, serious noise and color distortion. The main challenges of underwater image enhancement are to preserve details in dark regions while avoiding oversaturetion in bright regions. This paper proposes a novel underwater image enhancement method based on image decomposition. By decomposing the high-frequency texture and noise into the texture layer, the transmission map is estimated from the noise-free structure layer to avoid the noise amplification problem in underwater image enhancement. Both the structure layer and texture layer are descattered with the estimated transmission map. After denoising by gradient residual minimizition, the texture layer is enhanced and added back into the structure layer to recover the final enhanced image. Experimental results verify that the proposed approach can recover the high-quality images with fine details and edges while improving contrast and color naturalness, especially for images taken in the high turbidity environment. Jing-Yu Yang 0002, Huanjing Yue, Xiaomei Fu, Chunping Hou |
ICIP | 1 |
| 2017 | Low-rank matrix completion against missing rows and columns with separable 2-D sparsity priorsabstractMost existing matrix completion approaches assume that entries of matrices are missing at random, which could be violated in practical applications. This paper proposes a novel matrix completion method equipped with Joint Priors of LOw-rank and Separable 2-D Sparsity (JPLOSS) to complete missing rows and columns besides random missing. The underlying matrix is regularized by a low-rank prior, and its rows and columns are regularized by a row and a column dictionary, respectively. An reweighting scheme is incorporated into both the low-rank and sparsity terms to promote the low-rankness and sparseness simultaneously. The proposed model is effectively solved by an alternating direction method under the augmented Lagrangian multiplier framework. Experiments on both synthetic data and real images demonstrate the effectiveness and superiority of the proposed model in completing matrices with missing rows and columns compared with state-of-the-art matrix completion approaches. Jiaoru Yang, Kun Li 0001, Jing-Yu Yang 0002 |
ICIP | 4 |
| 2017 | Image noise estimation and removal considering the bayer pattern of noise varianceabstractTraditional image denoising methods are designed for Gaussian or Poisson noise, which are not suitable for realistic noise introduced in the complicated imaging pipeline. We observe that, due to the demosaicing process in imaging, the noise variance maps of captured JPEG images are characterized by Bayer patterns. In this paper, we propose a novel noise estimation and removal method based on the Bayer pattern of noise variance maps. There are two key contributions in the proposed method. First, to the best of our knowledge, we are the first to consider the Bayer patterns of noise variance maps in noise estimation and denoising. Second, we extend the state-of-the-art denoising method CBM3D to deal with realistic noise by integrating the estimated noise variance map and Bayer-pattern down-sampling into the denoising process. Experimental results show that the proposed method achieves the best noise estimation performance compared with two state-of-the-art methods. In addition, the denoising performance of CBM3D for realistic noise is significantly improved using the proposed approach and outperforms state-of-the-art blind denoising methods. Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Chunping Hou |
ICIP | 3 |
| 2017 | Global alignment of deformable objects captured by a single RGB-D cameraabstractWe present a novel global registration method for deformable objects captured using a single RGB-D camera. Our algorithm allows objects to undergo large non-rigid deformations, and achieves high quality results without constraining the actor's pose or camera motion. We compute the deformations of all the scans simultaneously by optimizing a global alignment problem to avoid the well-known loop closure problem, and use an as-rigid-as-possible constraint to eliminate the shrinkage problem of the deformed model. To attack large scale problems, we design a coarse-to-fine multi-resolution scheme, which also avoids the optimization being trapped into local minima. The proposed method is evaluated on public datasets and real datasets captured by an RGB-D sensor. Experimental results demonstrate that the proposed method obtains better results than the state-of-the-art methods. Daoliang Guo, Kun Li 0001, Yukun Lai, Jing-Yu Yang 0002 |
ICME | 4 |
| 2017 | 3-D motion recovery via low rank matrix restoration on articulation graphsabstractThis paper addresses the challenge of 3-D skeleton recovery by exploiting the spatio-temporal correlations of corrupted 3D skeleton sequences. A skeleton sequence is represented as a matrix. We propose a novel low-rank solution that effectively integrates both a low-rank model for robust skeleton recovery based on temporal coherence, and an articulation-graph-based isometric constraint for spatial coherence, namely consistency of bone lengths. The proposed model is formulated as a constrained optimization problem, which is efficiently solved by the Augmented Lagrangian Method with a Gauss-Newton solver for the subproblem of isometric optimization. Experimental results on the CMU motion capture dataset and a Kinect dataset show that the proposed approach achieves better recovery accuracy over a state-of-the-art method. The proposed method has wide applicability for skeleton tracking devices, such as the Kinect, because these devices cannot provide accurate reconstructions of complex motions, especially in the presence of occlusion. Kun Li 0001, Meiyuan Wang, Yukun Lai, Jing-Yu Yang 0002, Feng Wu 0001 |
ICME | 4 |
| 2017 | Intrinsic decomposition from a single RGB-D image with sparse and non-local priorsabstractThis paper proposes a new intrinsic image decomposition method that decomposes a single RGB-D image into reflectance and shading components. We observe and verify that, a shading image mainly contains smooth regions separated by curves, and its gradient distribution is sparse. We therefore use ℓ1-norm to model the direct irradiance component - the main sub-component extracted from shading component. Moreover, a non-local prior weighted by a bilateral kernel on a larger neighborhood is designed to fully exploit structural correlation in the reflectance component to improve the decomposition performance. The model is solved by the alternating direction method under the augmented Lagrangian multiplier (ADM-ALM) framework. Experimental results on both synthetic and real datasets demonstrate that the proposed method yields better results and enjoys lower complexity compared with two state-of-the-art methods. Kun Li 0001, Jing-Yu Yang 0002, Xinchen Ye |
ICME | 3 |
| 2017 | Depth super-resolution via fully edge-augmented guidanceabstractRecently, convolutional neural networks (CNNs) have been widely used for image processing problems. In this work, we present an end-to-end depth map super-resolution method based on CNN. Standing on a residual learning architecture, the proposed network learns joint features to get a high-resolution (HR) depth map from a low-resolution (LR) one with the multi-layers guidance of a HR color image. Furthermore, in order to focus on the boundaries of depth map, we generate an edge-attention map from the associated HR color images as a guidance. Experimental results show that the proposed network outperforms the state-of-the-art depth map super-resolution methods. Jing-Yu Yang 0002, Kun Li 0001 |
VCIP | 1 |
| 2017 | Demoiréing for screen-shot images with multi-channel layer decompositionabstractMoiré patterns on screen-shot images are mainly due to the aliasing of the grid of the display and the camera sensor, which heavily degenerated the image quality. This paper proposes an demoiréing method for screen-shot images via layer decomposition on polyphase components (LDPC). The layer decomposition model separates the image into a background layer and a moiré layer, which are both regularized by a patch-based Gaussian Mixture Model (GMM) prior. To enhance the distinguishability between the image patches and moiré patches, the input image is first subsampled into four polyphase components, each of which is decomposed with the GMM-based layer decomposition model. The proposed model is applied on luminance (Y) channel to weaken the intensity of moiré patterns, and on red, green, blue channels respectively to further remove moiré patterns. Experimental results demonstrate that the proposed method is able to efficiently remove moiré artifacts for screen-shot images and outperform several other methods. Jing-Yu Yang 0002, Changrui Cai, Kun Li 0001 |
VCIP | 1 |
| 2017 | SPA: Sparse Photorealistic Animation Using a Single RGB-D CameraabstractPhotorealistic animation is a desirable technique for computer games and movie production. We propose a new method to synthesize plausible videos of human actors with new motions using a single cheap RGB-D camera. A small database is captured in a usual office environment, which happens only once for synthesizing different motions. We propose a marker-less performance capture method using sparse deformation to obtain the geometry and pose of the actor for each time instance in the database. Then, we synthesize an animation video of the actor performing the new motion that is defined by the user. An adaptive model-guided texture synthesis method based on weighted low-rank matrix completion is proposed to be less sensitive to noise and outliers, which enables us to easily create photorealistic animation videos with new motions that are different from the motions in the database. Experimental results on the public data set and our captured data set have verified the effectiveness of the proposed method. Kun Li 0001, Jing-Yu Yang 0002, Leijie Liu, Ronan Boulic, Yukun Lai, Yebin Liu, Eray Molla |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Textured Image Demoiréing via Signal Decomposition and Guided FilteringabstractMoiré artifacts are generally caused by the interference between the overlap of the sensor's sampling grid and high-frequency (nearly) periodic textures, and heavily affect the image quality. However, it is difficult to effectively remove moiré artifacts from textured images as the structure of moiré patterns is similar to that of textures in some sense. In this paper, we propose a novel textured image demoiréing method by signal decomposition and guided filtering. Given a textured image with moiré artifacts, we first remove moiré artifacts in the green (G) channel using the proposed low-rank and sparse matrix decomposition model. This model regularizes the texture layer by the low-rank prior in spatial domain and the moiré layer by sparse representation in frequency domain. An alternating direction method under the augmented Lagrangian multiplier framework is used to solve the matrix decomposition model. Then, since the red (R) and blue (B) channels are more heavily polluted by moiré artifacts than the G channel, we propose to remove moiré artifacts in its R and B channels via guided filtering by the obtained texture layer of the G channel. Experimental results demonstrate that our method outperforms the state-of-the-art methods for both synthetic and real images. Jing-Yu Yang 0002, Fanglei Liu, Huanjing Yue, Xiaomei Fu, Chunping Hou, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Reconstruction of Structurally-Incomplete Matrices With Reweighted Low-Rank and Sparsity PriorsabstractMost matrix reconstruction methods assume that missing entries randomly distribute in the incomplete matrix, and the low-rank prior or its variants are used to well pose the problem. However, in practical applications, missing entries are structurally rather than randomly distributed, and cannot be handled by the rank minimization prior individually. To remedy this, this paper introduces new matrix reconstruction models using double priors on the latent matrix, named Reweighted Low-rank and Sparsity Priors (ReLaSP). In the proposed ReLaSP models, the matrix is regularized by a low-rank prior to exploit the inter-column and inter-row correlations, and its columns (rows) are regularized by a sparsity prior under a dictionary to exploit intra-column (-row) correlations. Both the low-rank and sparse priors are reweighted on the fly to promote low-rankness and sparsity, respectively. Numerical algorithms to solve our ReLaSP models are derived via the alternating direction method under the augmented Lagrangian multiplier framework. Results on synthetic data, image restoration tasks, and seismic data interpolation show that the proposed ReLaSP models are quite effective in recovering matrices degraded by highly structural missing and various types of noise, complementing the classic matrix reconstruction models that handle random missing only. Jing-Yu Yang 0002, Xuemeng Yang, Xinchen Ye, Chunping Hou |
IEEE Trans. Image Process. | 1 |
| 2017 | Contrast Enhancement Based on Intrinsic Image DecompositionabstractIn this paper, we propose to introduce intrinsic image decomposition priors into decomposition models for contrast enhancement. Since image decomposition is a highly illposed problem, we introduce constraints on both reflectance and illumination layers to yield a highly reliable solution. We regularize the reflectance layer to be piecewise constant by introducing a weighted ℓ1norm constraint on neighboring pixels according to the color similarity, so that the decomposed reflectance would not be affected much by the illumination information. The illumination layer is regularized by a piecewise smoothness constraint. The proposed model is effectively solved by the Split Bregman algorithm. Then, by adjusting the illumination layer, we obtain the enhancement result. To avoid potential color artifacts introduced by illumination adjusting and reduce computing complexity, the proposed decomposition model is performed on the value channel in HSV space. Experiment results demonstrate that the proposed method performs well for a wide variety of images, and achieves better or comparable subjective and objective quality compared with the state-of-the-art methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001, Chunping Hou |
IEEE Trans. Image Process. | 2 |
| 2016 | Completion of structurally-incomplete matrices with reweighted low-rank and sparsity priorsabstractMost matrix completion methods impose a low-rank prior or its variants to well pose the problem. However, the rank minimization is problematic to handle matrices with structural missing. To remedy this, this paper introduces a new matrix completion method using double priors on the latent matrix, named Reweighted Low-rank and Sparsity Priors. In the proposed model, the matrix is regularized by a low-rank prior to exploit the inter-column (row) correlations, and its columns (rows) are regularized by a sparsity prior under a dictionary to exploit intra-column (row) correlations. Both the low-rank and sparse priors are reweighted on the fly to promote low-rankness and sparsity, respectively. Numerical algorithm to solve our model is derived via the alternating direction method under the augmented Lagrangian multiplier framework. Experimental results show that our model is quite effective in recovering matrices with highly-structural missing, complementing the classic matrix completion models that handle random missing only. Jing-Yu Yang 0002, Xuemeng Yang, Xinchen Ye |
ICASSP | 1 |
| 2016 | Depth refinement for binocular kinect RGB-D camerasabstractThis paper presents a novel depth refinement framework for binocular Kinect RGB-D cameras for obtaining high quality depth map. Firstly, we build a binocular depth sensing system using two Kinect v2 cameras, and analyze the systematic error of the system from two aspects, i.e., camera interaction and intrinsic characteristics. Then, the captured depth maps from different views are fused to fully exploit the inter-view correlations, and an error compensation method is proposed to remove the systematic errors from the fused depth map. Finally, an edge-guided depth propagation scheme is used to refine the depth map from binocular depth map. Experimental results show that the proposed framework is able to substantially improve the quality of depth image. Jinghui Bai, Jing-Yu Yang 0002, Xinchen Ye, Chunping Hou |
VCIP | 2 |
| 2016 | 3-D motion recovery via low rank matrix analysisabstractSkeleton tracking is a useful and popular application of Kinect. However, it cannot provide accurate reconstructions for complex motions, especially in the presence of occlusion. This paper proposes a new 3-D motion recovery method based on low-rank matrix analysis to correct invalid or corrupted motions. We address this problem by representing a motion sequence as a matrix, and introducing a convex low-rank matrix recovery model, which fixes erroneous entries and finds the correct low-rank matrix by minimizing nuclear norm and norm of constituent clean motion and error matrices. Experimental results show that our method recovers the corrupted skeleton joints, achieving accurate and smooth reconstructions even for complicated motions. Meiyuan Wang, Kun Li 0001, Feng Wu 0001, Yukun Lai, Jing-Yu Yang 0002 |
VCIP | 5 |
| 2016 | Background recovery from video sequences via online motion-assisted RPCAabstractBackground modeling is an important technique for video analysis. Robust principal component analysis (RPCA) assisted with motion information has shown improved background recovery performance, but still suffers from the deficiency in handling steaming video due to the batch-mode formulation and implementation. This paper proposes an online motion-assisted robust principal component analysis (OMA-RPCA) model for background recovery from video sequences. The inherent batch-mode nuclear norm for low-rank approximation is replaced with an explicitly low-rank matrix factorization. Motion information extracted by an optical flow method is incorporated into the data term to facilitate the separation of moving objects from the background. The proposed model is effectively solved by an alternating optimization scheme in an online mode. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods with lower memory cost and scalability to online applications. Jiaoru Yang, Jing-Yu Yang 0002, Xuemeng Yang, Huanjing Yue |
VCIP | 2 |
| 2016 | Depth recovery via decomposition of polynomial and piece-wise constant signalsabstractThis paper proposes a novel decomposition model for high-quality depth recovery (DMDR) from low quality depth measurement accompanied by high-resolution RGB image. We observe that depth patches extracted from the depth map containing smooth regions separated by curves, can be decomposed simultaneously by a low-order polynomial surface and a piece-wise constant signal. In our model, the polynomial surface component is regularized by least-square polynomial smoothing, while the piece-wise constant component is constrained by total variation filtering. The model is effectively solved by the alternating direction method under the augmented Lagrangian multiplier (ALM-ADM) algorithm. Experimental results show that our method is able to handle various types of depth degradation under the designed signal decomposition model, and produces high-quality depth recovery results. Xinchen Ye, Jing-Yu Yang 0002, Chunping Hou, Yao Wang 0001 |
VCIP | 3 |
| 2016 | Video super-resolution using an adaptive superpixel-guided auto-regressive model
Kun Li 0001, Yanming Zhu 0001, Jing-Yu Yang 0002, Jianmin Jiang |
Pattern Recognit. | 3 |
| 2016 | Lossless Compression of JPEG Coded Photo CollectionsabstractThe explosion of digital photos has posed a significant challenge to photo storage and transmission for both personal devices and cloud platforms. In this paper, we propose a novel lossless compression method to further reduce the size of a set of JPEG coded correlated images without any loss of information. The proposed method jointly removes inter/intra image redundancy in the feature, spatial, and frequency domains. For each collection, we first organize the images into a pseudo video by minimizing the global prediction cost in the feature domain. We then present a hybrid disparity compensation method to better exploit both the global and local correlations among the images in the spatial domain. Furthermore, the redundancy between each compensated signal and the corresponding target image is adaptively reduced in the frequency domain. Experimental results demonstrate the effectiveness of the proposed lossless compression method. Compared with the JPEG coded image collections, our method achieves average bit savings of more than 31%. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Rectangle fitting via quadratic programmingabstractThis paper investigates rectangle fitting via optimization approaches. We summarize two basic requirements for rectangular fitting, leading to a basic model that are non-convex and difficult to attack. To avoid potential trapping of local minima, we extend the basic model with centroid and orientation constraints into a quadratic programming. To achieve reliable fitting from noisy points, slack variables are introduced to soften hard constraints. The scalability to problem size are further addressed by careful selecting only a small fraction of slack variables. Results on clean dataset, noisy dataset, and practical data show that our method is able to reliably fit rectangles for various kinds of data. Jing-Yu Yang 0002, Zhongyu Jiang |
MMSP | 1 |
| 2015 | RIFO: Restoring images with fence occlusionsabstractMany scenes, e.g., zoos, parks, and gardens, are guarded by fences, and people can only take pictures through the fences. It is desirable to remove visually-annoying fence occlusions from images. This paper proposes a novel approach to restore images from fence occlusions (RIFO). The proposed method consists of two steps: fence detection, and disocclusion restoration. In fence detection, the image is first clustered into superpixels, which are fitted into rectangles. We collect fence pixels from superpixels by determining the elongation of their fitted rectangles. A primary shape of the fence is obtained by a color-based classifier learned from sampled pixels. Then, multi-RANSAC and moving least squares (MLS) are used for sketching the fence structure. Complete fence is detected by expanding the fence structure. Disoccluded regions are restorated by a patch-based approach using matrix completion. Experimental results show that our method detects complete fences from images, and the disoccluded regions are faithfully recovered, yielding clean and complete images. Jing-Yu Yang 0002, Leijie Liu, Chunping Hou |
MMSP | 1 |
| 2015 | Moiré pattern removal from texture images via low-rank and sparse matrix decompositionabstractMoiré patterns, an artifact of aliasing interference between details in the subject matter and the grid of the sensor, heavily disturb the qualitative and quantitative analysis of images. It is hard to effectively remove moiré patterns since they are similar to image textures. We propose a novel low-rank and sparse matrix decomposition model for moiré pattern removal. This method is grounded on the observation: textures are locally well-patterned while moiré patterns are dissimilar, and the energy distribution of moiré patterns in the frequency domain is concentrated and almost no mixed with that of textures. For each patch, texture component is regularized by a low-rank prior and moiré component is regularized by a sparse prior in the discrete cosine transform (DCT) domain. This model is effectively solved by the alternating direction method under the augmented Lagrangian multiplier (ALM-ADM) algorithm. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods. Fanglei Liu, Jing-Yu Yang 0002, Huanjing Yue |
VCIP | 2 |
| 2015 | Incremental SfM based lossless compression of JPEG coded photo albumabstractThe key problem in photo album compression is how to exploit the correlation among the images. In this paper, we propose a novel incremental structure from motion (SfM) based prediction method for lossless photo album compression. Unlike the previous methods, we exploit the redundancy among images through their inherent geometric relationship generated by SfM. Based on the point cloud and camera poses, each prediction image is generated by projecting, triangulation and warping. Finally, the target image is compressed by an HEVC-like encoder with the prediction image as main reference. Experimental results demonstrate the advantage of our method, especially for images with scenes containing complicated geometric structures. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
VCIP | 3 |
| 2015 | Sparse Non-rigid Registration of 3D ShapesabstractAbstract Non‐rigid registration of 3D shapes is an essential task of increasing importance as commodity depth sensors become more widely available for scanning dynamic scenes. Non‐rigid registration is much more challenging than rigid registration as it estimates a set of local transformations instead of a single global transformation, and hence is prone to the overfitting issue due to underdetermination. The common wisdom in previous methods is to impose an ℓ2‐norm regularization on the local transformation differences. However, the ℓ2‐norm regularization tends to bias the solution towards outliers and noise with heavy‐tailed distribution, which is verified by the poor goodness‐of‐fit of the Gaussian distribution over transformation differences. On the contrary, Laplacian distribution fits well with the transformation differences, suggesting the use of a sparsity prior. We propose a sparse non‐rigid registration (SNR) method with an ℓ1‐norm regularized model for transformation estimation, which is effectively solved by an alternate direction method (ADM) under the augmented Lagrangian framework. We also devise a multi‐resolution scheme for robust and progressive registration. Results on both public datasets and our scanned datasets show the superiority of our method, particularly in handling large‐scale deformations as well as outliers and noise. Jing-Yu Yang 0002, Kun Li 0001, Yukun Lai |
Comput. Graph. Forum | 1 |
| 2015 | Foreground-Background Separation From Video Clips via Motion-Assisted Matrix RestorationabstractSeparation of video clips into foreground and background components is a useful and important technique, making recognition, classification, and scene analysis more efficient. In this paper, we propose a motion-assisted matrix restoration (MAMR) model for foreground-background separation in video clips. In the proposed MAMR model, the backgrounds across frames are modeled by a low-rank matrix, while the foreground objects are modeled by a sparse matrix. To facilitate efficient foreground-background separation, a dense motion field is estimated for each frame, and mapped into a weighting matrix which indicates the likelihood that each pixel belongs to the background. Anchor frames are selected in the dense motion estimation to overcome the difficulty of detecting slowly moving objects and camouflages. In addition, we extend our model to a robust MAMR model against noise for practical applications. Evaluations on challenging datasets demonstrate that our method outperforms many other state-of-the-art methods, and is versatile for a wide range of surveillance videos. Xinchen Ye, Jing-Yu Yang 0002, Kun Li 0001, Chunping Hou, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Nonrigid Structure From Motion via Sparse RepresentationabstractThis paper proposes a new approach for nonrigid structure from motion with occlusion, based on sparse representation. We address the occlusion problem based on the latest developments on sparse representation: matrix completion, which can recover the observation matrix that has high percentages of missing data and can also reduce the noises and outliers in the known elements. We introduce sparse transform to the joint estimation of 3-D shapes and motions. 3-D shape trajectory space is fit by wavelet basis to achieve better modeling of complex motion. Experimental results on datasets without and with occlusion show that our method can better estimate the 3-D shapes and motions, compared with state-of-the-art algorithms. Kun Li 0001, Jing-Yu Yang 0002, Jianmin Jiang |
IEEE Trans. Cybern. | 2 |
| 2015 | Graph-Based Segmentation for RGB-D Data Using 3-D Geometry Enhanced SuperpixelsabstractWith the advances of depth sensing technologies, color image plus depth information (referred to as RGB-D data hereafter) is more and more popular for comprehensive description of 3-D scenes. This paper proposes a two-stage segmentation method for RGB-D data: 1) oversegmentation by 3-D geometry enhanced superpixels and 2) graph-based merging with label cost from superpixels. In the oversegmentation stage, 3-D geometrical information is reconstructed from the depth map. Then, a K-means-like clustering method is applied to the RGB-D data for oversegmentation using an 8-D distance metric constructed from both color and 3-D geometrical information. In the merging stage, treating each superpixel as a node, a graph-based model is set up to relabel the superpixels into semantically-coherent segments. In the graph-based model, RGB-D proximity, texture similarity, and boundary continuity are incorporated into the smoothness term to exploit the correlations of neighboring superpixels. To obtain a compact labeling, the label term is designed to penalize labels linking to similar superpixels that likely belong to the same object. Both the proposed 3-D geometry enhanced superpixel clustering method and the graph-based merging method from superpixels are evaluated by qualitative and quantitative results. By the fusion of color and depth information, the proposed method achieves superior segmentation performance over several state-of-the-art algorithms. Jing-Yu Yang 0002, Ziqiao Gan, Kun Li 0001, Chunping Hou |
IEEE Trans. Cybern. | 1 |
| 2015 | Estimation of Signal-Dependent Noise Level Function in Transform Domain via a Sparse Recovery ModelabstractThis paper proposes a novel algorithm to estimate the noise level function (NLF) of signal-dependent noise (SDN) from a single image based on the sparse representation of NLFs. Noise level samples are estimated from the high-frequency discrete cosine transform (DCT) coefficients of nonlocal-grouped low-variation image patches. Then, an NLF recovery model based on the sparse representation of NLFs under a trained basis is constructed to recover NLF from the incomplete noise level samples. Confidence levels of the NLF samples are incorporated into the proposed model to promote reliable samples and weaken unreliable ones. We investigate the behavior of the estimation performance with respect to the block size, sampling rate, and confidence weighting. Simulation results on synthetic noisy images show that our method outperforms existing state-of-the-art schemes. The proposed method is evaluated on real noisy images captured by three types of commodity imaging devices, and shows consistently excellent SDN estimation performance. The estimated NLFs are incorporated into two well-known denoising schemes, nonlocal means and BM3D, and show significant improvements in denoising SDN-polluted images. Jing-Yu Yang 0002, Ziqiao Gan, Zhaoyang Wu, Chunping Hou |
IEEE Trans. Image Process. | 1 |
| 2015 | Image Denoising by Exploring External and Internal CorrelationsabstractSingle image denoising suffers from limited data collection within a noisy image. In this paper, we propose a novel image denoising scheme, which explores both internal and external correlations with the help of web images. For each noisy patch, we build internal and external data cubes by finding similar patches from the noisy and web images, respectively. We then propose reducing noise by a two-stage strategy using different filtering approaches. In the first stage, since the noisy patch may lead to inaccurate patch selection, we propose a graph based optimization method to improve patch matching accuracy in external denoising. The internal denoising is frequency truncation on internal cubes. By combining the internal and external denoising patches, we obtain a preliminary denoising result. In the second stage, we propose reducing noise by filtering of external and internal cubes, respectively, on transform domain. In this stage, the preliminary denoising result not only enhances the patch matching accuracy but also provides reliable estimates of filtering parameters. The final denoising image is obtained by fusing the external and internal filtering results. Experimental results show that our method constantly outperforms state-of-the-art denoising schemes in both subjective and objective quality measurements, e.g., it achieves >2 dB gain compared with BM3D at a wide range of noise levels. Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | CID: Combined Image Denoising in Spatial and Frequency Domains Using Web ImagesabstractIn this paper, we propose a novel two-step scheme to filter heavy noise from images with the assistance of retrieved Web images. There are two key technical contributions in our scheme. First, for every noisy image block, we build two three dimensional (3D) data cubes by using similar blocks in retrieved Web images and similar nonlocal blocks within the noisy image, respectively. To better use their correlations, we propose different denoising strategies. The denoising in the 3D cube built upon the retrieved images is performed as median filtering in the spatial domain, whereas the denoising in the other 3D cube is performed in the frequency domain. These two denoising results are then combined in the frequency domain to produce a denoising image. Second, to handle heavy noise, we further propose using the denoising image to improve image registration of the retrieved Web images, 3D cube building, and the estimation of filtering parameters in the frequency domain. Afterwards, the proposed denoising is performed on the noisy image again to generate the final denoising result. Our experimental results show that when the noise is high, the proposed scheme is better than BM3D by more than 2 dB in PSNR and the visual quality improvement is clear to see. Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
CVPR | 3 |
| 2014 | Background extraction from video sequences via motion-assisted matrix completionabstractBackground extraction from video sequences is a useful and important technique in video surveillance. This paper proposes a motion-assisted matrix completion model for background extraction from video sequences. A binary motion map is first calculated for each frame by optical flow. By excluding areas associated with moving objects with the binary motion maps, the background extraction is formulated into a motion-assisted matrix completion (MAMC) problem. Experimental results show that our method not only extracts promising backgrounds but also outperforms many state-of-the-art methods in distinguishing moving objects on challenging datasets. Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001 |
ICIP | 1 |
| 2014 | Non-rigid structure from motion via sparse representationabstractThis paper proposes a new approach for non-rigid structure from motion with occlusion, based on sparse representation. We introduce sparse transform to the joint estimation of 3D shapes and motions. 3D shape trajectory space is fit by wavelet basis to achieve better modeling of complex motion. We address the occlusion problem based on the latest developments on sparse representation: matrix completion, which can recover the observation matrix that has high percentages of missing data and can also reduce the noises and outliers in the known elements. Experimental results on datasets without and with occlusion show that our method can better estimate the 3D shapes and motions, compared with state-of-the-art algorithms. Kun Li 0001, Jing-Yu Yang 0002, Jianmin Jiang |
ICME | 2 |
| 2014 | Lossless compression of JPEG coded photo albumsabstractThe explosion in digital photography poses a significant challenge when it comes to photo storage for both personal devices and the Internet. In this paper, we propose a novel lossless compression method to further reduce the storage size of a set of JPEG coded correlated images. In this method, we propose jointly removing the inter-image redundancy in the feature, spatial, and frequency domains. For each album, we first organize the images into a pseudo video by minimizing the global predictive cost in the feature domain. We then introduce a disparity compensation method to enhance the spatial correlation between images. Finally, the redundancy between the compensated signal and the corresponding target image is adaptively reduced in the frequency domain. Moreover, our proposed scheme is able to losslessly recover not only raw images but also JPEG files. Experimental results demonstrate the efficiency of our proposed lossless compression, which achieves more than 12% bit-saving on average compared with JPEG coded albums. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
VCIP | 3 |
| 2014 | Color-Guided Depth Recovery From RGB-D Data Using an Adaptive Autoregressive ModelabstractThis paper proposes an adaptive color-guided autoregressive (AR) model for high quality depth recovery from low quality measurements captured by depth cameras. We observe and verify that the AR model tightly fits depth maps of generic scenes. The depth recovery task is formulated into a minimization of AR prediction errors subject to measurement consistency. The AR predictor for each pixel is constructed according to both the local correlation in the initial depth map and the nonlocal similarity in the accompanied high quality color image. We analyze the stability of our method from a linear system point of view, and design a parameter adaptation scheme to achieve stable and accurate depth recovery. Quantitative and qualitative evaluation compared with ten state-of-the-art schemes show the effectiveness and superiority of our method. Being able to handle various types of depth degradations, the proposed method is versatile for mainstream depth sensors, time-of-flight camera, and Kinect, as demonstrated by experiments on real systems. Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou, Yao Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | SIFT-based image super-resolutionabstractThis paper presents a new exemplar-based image super-resolution (SR) method in which we propose making use of scale invariant image features for high frequency (HF) approximation. We introduce the scale invariant feature transform (SIFT) descriptors in both building an exemplar dataset adaptively and producing the HF details with respect to the features of an input low resolution image. Given a large image database, we propose using the highly correlated images retrieved by SIFT descriptors for exemplar training rather than using a general set of images to increase the matching accuracy. Through building the training set of high resolution/low resolution exemplar pairs, the HF details for SR are retrieved from the training set by matching the SIFT features in a dense way. The flexibility as well as effectiveness of our SR approach is demonstrated at different magnification factors, e.g. 3 and 4. Experimental results show that our SIFT-based SR approach achieves enhanced high resolution images in terms of both objective and subjective qualities in comparison with the state-of-the-art exemplar-based methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
ISCAS | 2 |
| 2013 | Landmark Image Super-Resolution by Retrieving Web ImagesabstractThis paper proposes a new super-resolution (SR) scheme for landmark images by retrieving correlated web images. Using correlated web images significantly improves the exemplar-based SR. Given a low-resolution (LR) image, we extract local descriptors from its up-sampled version and bundle the descriptors according to their spatial relationship to retrieve correlated high-resolution (HR) images from the web. Though similar in content, the retrieved images are usually taken with different illumination, focal lengths, and shot perspectives, resulting in uncertainty for the HR detail approximation. To solve this problem, we first propose aligning these images to the up-sampled LR image through a global registration, which identifies the corresponding regions in these images and reduces the mismatching. Second, we propose a structure-aware matching criterion and adaptive block sizes to improve the mapping accuracy between LR and HR patches. Finally, these matched HR patches are blended together by solving an energy minimization problem to recover the desired HR image. Experimental results demonstrate that our SR scheme achieves significant improvement compared with four state-of-the-art schemes in terms of both subjective and objective qualities. Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Cloud-Based Image Coding for Mobile Devices - Toward Thousands to One CompressionabstractCurrent image coding schemes make it hard to utilize external images for compression even if highly correlated images can be found in the cloud. To solve this problem, we propose a method of cloud-based image coding that is different from current image coding even on the ground. It no longer compresses images pixel by pixel and instead tries to describe images and reconstruct them from a large-scale image database via the descriptions. First, we describe an input image based on its down-sampled version and local feature descriptors. The descriptors are used to retrieve highly correlated images in the cloud and identify corresponding patches. The down-sampled image serves as a target to stitch retrieved image patches together. Second, the down-sampled image is compressed using current image coding. The feature vectors of local descriptors are predicted by the corresponding vectors extracted in the decoded down-sampled image. The predicted residual vectors are compressed by transform, quantization, and entropy coding. The experimental results show that the visual quality of reconstructed images is significantly better than that of intra-frame coding in HEVC and JPEG at thousands to one compression . Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2012 | Depth Recovery Using an Adaptive Color-Guided Auto-Regressive Model
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou |
ECCV (5) | 1 |
| 2012 | Estimation of signal-dependent sensor noise via sparse representation of noise level functionsabstractThis paper proposes a noise estimation method for signal-dependent sensor noise based on sparse representation of noise level functions (NLFs). Homogeneous blocks are detected by an image structure analyzer, and grouped to estimate noise levels for various image intensities with confidences. The noise level function is recovered from the incomplete and noisy estimated samples by solving its sparse representation under a trained basis. Experimental results show that our proposed method accurately estimates NLFs for both smooth and highly-textured images over various noise levels. Jing-Yu Yang 0002, Zhaoyang Wu, Chunping Hou |
ICIP | 1 |
| 2012 | SIFT-Based Image CompressionabstractThis paper proposes a novel image compression scheme based on the local feature descriptor - Scale Invariant Feature Transform (SIFT). The SIFT descriptor characterizes an image region invariantly to scale and rotation. It is used widely in image retrieval. By using SIFT descriptors, our compression scheme is able to make use of external image contents to reduce visual redundancy among images. The proposed encoder compresses an input image by SIFT descriptors rather than pixel values. It separates the SIFT descriptors of the image into two groups, a visual description which is a significantly sub sampled image with key SIFT descriptors embedded and a set of differential SIFT descriptors, to reduce the coding bits. The corresponding decoder generates the SIFT descriptors from the visual description and the differential set. The SIFT descriptors are used in our SIFT-based matching to retrieve the candidate predictive patches from a large image dataset. These candidate patches are then integrated into the visual description, presenting the final reconstructed images. Our preliminary but promising results demonstrate the effectiveness of our proposed image coding scheme towards perceptual quality. Our proposed image compression scheme provides a feasible approach to make use of the visual correlation among images. Huanjing Yue, Xiaoyan Sun 0001, Feng Wu 0001, Jing-Yu Yang 0002 |
ICME | 4 |
| 2012 | Three-Dimensional Motion Estimation via Matrix CompletionabstractThree-dimensional motion estimation from multiview video sequences is of vital importance to achieve high-quality dynamic scene reconstruction. In this paper, we propose a new 3-D motion estimation method based on matrix completion. Taking a reconstructed 3-D mesh as the underlying scene representation, this method automatically estimates motions of 3-D objects. A "separating + merging" framework is introduced to multiview 3-D motion estimation. In the separating step, initial motions are first estimated for each view with a neighboring view. Then, in the merging step, the motions obtained by each view are merged together and optimized by low-rank matrix completion method. The most accurate motion estimation for each vertex in the recovered matrix is further selected by three spatiotemporal criteria. Experimental results on data sets with synthetic motions and real motions show that our method can reliably estimate 3-D motions. Kun Li 0001, Qionghai Dai, Wenli Xu, Jing-Yu Yang 0002, Jianmin Jiang |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2010 | Image coding via sparse contourlet representationabstractWhen contourlet coefficients are directly coded, the benefits from directional representation are discounted by the redundancy of the transform. In this paper, we propose a contourlet-based image coding scheme under the sparse representation framework, in which contourlet coefficients are sparsified by the iterative thresholding method before compression. Dependency analysis is performed to reveal correlations among sparsified contourlet coefficients. Considering the characteristics of coefficient correlations, the obtained coefficients are coded with an intra-subband coder. Experimental results show that the coding performances of the contourlet-based schemes are significantly boosted via sparse representations. Jing-Yu Yang 0002, Chunping Hou, Wenli Xu |
ISCAS | 1 |
| 2010 | Fast adaptive wavelet packets using interscale embedding of decomposition structures
Jing-Yu Yang 0002, Wenli Xu, Qionghai Dai |
Pattern Recognit. Lett. | 1 |
| 2009 | Image fusion in compressed sensingabstractThis paper proposes an efficient image fusion scheme for compressed sensing (CS) imaging, in which fusion is performed on the random projections before reconstruction. Specifically, the measurements of multiple input images are fused into composite measurements via weighted average, in which the weights are calculated based on entropy metrics of the original measurements. Then the fused image with transformation coefficients in a selected basis is reconstructed from the composite measurements by the gradient projection for sparse reconstruction (GPSR) algorithm. The proposed scheme is implemented in a block-based CS framework. Simulation results show that our scheme provides promising fusion performance with a low computational complexity. Xiaoyan Luo, Jun Zhang 0007, Jing-Yu Yang 0002, Qionghai Dai |
ICIP | 3 |
| 2009 | Ways to sparse representation: An overview
Jing-Yu Yang 0002, YiGang Peng, Wenli Xu, Qionghai Dai |
Sci. China Ser. F Inf. Sci. | 1 |
| 2009 | Image and Video Denoising Using Adaptive Dual-Tree Discrete Wavelet PacketsabstractWe investigate image and video denoising using adaptive dual-tree discrete wavelet packets (ADDWP), which is extended from the dual-tree discrete wavelet transform (DDWT). With ADDWP, DDWT subbands are further decomposed into wavelet packets with anisotropic decomposition, so that the resulting wavelets have elongated support regions and more orientations than DDWT wavelets. To determine the decomposition structure, we develop a greedy basis selection algorithm for ADDWP, which has significantly lower computational complexity than a previously developed optimal basis selection algorithm, with only slight performance loss. For denoising the ADDWP coefficients, a statistical model is used to exploit the dependency between the real and imaginary parts of the coefficients. The proposed denoising scheme gives better performance than several state-of-the-art DDWT-based schemes for images with rich directional features. Moreover, our scheme shows promising results without using motion estimation in video denoising. The visual quality of images and videos denoised by the proposed scheme is also superior. Jing-Yu Yang 0002, Yao Wang 0001, Wenli Xu, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | 2-D anisotropic dual-tree complex wavelet packets and its application to image denoisingabstractIn this paper, we extend the 2-D dual-tree complex wavelet transform (DTCWT) to an adaptive anisotropic dual-tree complex wavelet packets (ADTCWP). The DTCWT subbands are iteratively decomposed into anisotropic complex wavelet packets, generating anisotropic wavelets and increasing the number of wavelets orientations without introducing extra redundancy. Then a basis selection procedure is applied so that the selected complex wavelet packets are well adapted to image characteristics. The effectiveness of ADTCWP is examined in image denoising with a bivariate statistical model. The ADTCWP-based denoising scheme shows better denoising results than several DTCWT-based methods for image with rich directional features. The denoised images recovered by the proposed scheme are visually more appealing. Jing-Yu Yang 0002, Wenli Xu, Yao Wang 0001, Qionghai Dai |
ICIP | 1 |
| 2008 | Image Coding Using Dual-Tree Discrete Wavelet TransformabstractIn this paper, we explore the application of 2-D dual-tree discrete wavelet transform (DDWT), which is a directional and redundant transform, for image coding. Three methods for sparsifying DDWT coefficients, i.e., matching pursuit, basis pursuit, and noise shaping, are compared. We found that noise shaping achieves the best nonlinear approximation efficiency with the lowest computational complexity. The interscale, intersubband, and intrasubband dependency among the DDWT coefficients are analyzed. Three subband coding methods, i.e., SPIHT, EBCOT, and TCE, are evaluated for coding DDWT coefficients. Experimental results show that TCE has the best performance. In spite of the redundancy of the transform, our DDWT _ TCE scheme outperforms JPEG2000 up to 0.70 dB at low bit rates and is comparable to JPEG2000 at high bit rates. The DDWT _TCE scheme also outperforms two other image coders that are based on directional filter banks. To further improve coding efficiency, we extend the DDWT to an anisotropic dual-tree discrete wavelet packets (ADDWP), which incorporates adaptive and anisotropic decomposition into DDWT. The ADDWP subbands are coded with TCE coder. Experimental results show that ADDWP _ TCE provides up to 1.47 dB improvement over the DDWT _TCE scheme, outperforming JPEG2000 up to 2.00 dB. Reconstructed images of our coding schemes are visually more appealing compared with DWT-based coding schemes thanks to the directionality of wavelets. Jing-Yu Yang 0002, Yao Wang 0001, Wenli Xu, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2007 | Image Coding using 2-D Anisotropic Dual-Tree Discrete Wavelet TransformabstractWe propose an image coding scheme using 2-D anisotropic dual-tree discrete wavelet transform (DDWT). First, we extend 2-D DDWT to anisotropic decomposition, and obtain more directional subbands. Second, an iterative projection-based noise shaping algorithm is employed to further sparsify anisotropic DDWT coefficients. At last, the resulting coefficients are rearranged to preserve zero-tree relationship so that they can be efficiently coded with SPIHT. Experimental results show that our proposed scheme outperforms JPEG2000 and SPIHT at low bit rates despite the redundancy of DDWT. Jing-Yu Yang 0002, Jizheng Xu, Feng Wu 0001, Qionghai Dai, Yao Wang 0001 |
ICIP (3) | 1 |
| 2007 | Video Coding using 3-D Anisotropic Dual-Tree Wavelet TransformabstractThis paper investigates the use of the anisotropic 3-D dual-tree discrete wavelet transform (DDWT) for video coding. The 3-D DDWT is an attractive video representation because it isolates image patterns with different spatial orientations and motion directions and speeds in separate subbands. Our previous codecs using the 3-D isotropic DDWT provides better performance than the 3-D SPIHT codec on the 3-D DWT. In this paper, we explore the use of anisotropic DDWT (ADDWT) for video coding. The ADDWT extends the superiority of the normal DDWT with more directional subbands without adding to the redundancy. The proposed codec applies SPIHT to each of the ADDWT trees. This codec provides significantly better performance than the 3-D SPIHT codec using the standard DWT and the DDWT both objectively and subjectively. None of these video codecs requires motion compensation. Jing-Yu Yang 0002, Beibei Wang 0005, Yao Wang 0001, Wenli Xu |
ICME | 1 |
| 2007 | Image Compression using 2D Dual-tree Discrete Wavelet Transform (DDWT)abstractIn this paper, we investigate image compression using 2D dual-tree discrete wavelet transform (DDWT), which is an overcomplete transform with direction-selective basis functions. To further sparsify DDWT coefficients, an iterative projection-based noise shaping method is employed. We analyze the statistics of DDWT coefficients as well as the inter-scale, inter-subband, and intra-subband dependency among the DDWT coefficients. We further evaluate the application of SPIHT and EBCOT for coding DDWT coefficients. Experimental results show that SPIHT is more effective than EBCOT for DDWT, and the DDWT-SPIHT coder outperforms JPEG2000 at low bit rates and is comparable to JPEG2000 at high bit rates. Jing-Yu Yang 0002, Wenli Xu, Qionghai Dai, Yao Wang 0001 |
ISCAS | 1 |