EDBT 2026 Demo / reviewers in the wild / expert
Chen Zhao 0002
dblp:81/3-2
· DBLP profile ↗
46ranked-venue papers
17as first author
27since 2021 · last 2026
0000-0003-4993-5416ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 14 first-author · 20 since 2021Artificial intelligence and machine learning · 27 · 5 first-author · 26 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AddSR: Accelerating diffusion-based blind super-resolution with adversarial diffusion distillation
Ying Tai, Rui Xie 0005, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2026 | Spiking pyramid wavelet transformation for high-efficient and low-energy image restoration
Chen Zhao 0002, Xiantao Hu, Rui Xie 0005, Jian Yang 0003, Ying Tai |
Pattern Recognit. | 1 |
| 2026 | Learning multi-scale spatial-frequency features for image denoising
Xu Zhao 0001, Chen Zhao 0002, Xiantao Hu, Hongliang Zhang 0002, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 2 |
| 2025 | Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingabstractMultimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003 |
AAAI | 4 |
| 2025 | BOLT: Boost Large Vision-Language Model Without Training for Long-form Video UnderstandingabstractLarge video-language models (VLMs) have demonstrated promising progress in various video understanding tasks. However, their effectiveness in long-form video analysis is constrained by limited context windows. Traditional approaches, such as uniform frame sampling, often inevitably allocate resources to irrelevant content, diminishing their effectiveness in real-world scenarios. In this paper, we introduce BOLT, a method to BOost Large VLMs without additional Training through a comprehensive study of frame selection strategies. First, to enable a more realistic evaluation of VLMs in long-form video understanding, we propose a multi-source retrieval evaluation setting. Our findings reveal that uniform sampling performs poorly in noisy contexts, underscoring the importance of selecting the right frames. Second, we explore several frame selection strategies based on query-frame similarity and analyze their effectiveness at inference time. Our results show that inverse transform sampling yields the most significant performance improvement, increasing accuracy on the Video-Mme benchmark from 53.8% to 56.1% and MLVU benchmark from 58.9% to 63.4%. Our code is available at https://github.com/sming256/BOLT. Shuming Liu 0001, Chen Zhao 0002, Bernard Ghanem |
CVPR | 2 |
| 2025 | OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure CorrectionabstractExposure correction is a fundamental problem in computer vision and image processing. Recently, frequency domainbased methods have achieved impressive improvement, yet they still struggle with complex real-world scenarios under extreme exposure conditions. This is due to the local convolutional receptive fields failing to model long-range dependencies in the spectrum, and the non-generative learning paradigm being inadequate for retrieving lost details from severely degraded regions. In this paper, we propose Omnidirectional Spectral Mamba (OSMamba), a novel exposure correction network that incorporates the advantages of state space models and generative diffusion models to address these limitations. Specifically, OSMamba introduces an omnidirectional spectral scanning mechanism that adapts Mamba to the frequency domain to capture comprehensive long-range dependencies in both the amplitude and phase spectra of deep image features, hence enhancing illumination correction and structure recovery. Furthermore, we develop a dual-domain prior generator that learns from well-exposed images to generate a degradation-free diffusion prior containing correct information about severely under- and over-exposed regions for better detail restoration. Extensive experiments on multiple-exposure and mixed-exposure datasets demonstrate that the proposed OSMamba achieves state-of-the-art performance both quantitatively and qualitatively. Our code and models can be found at https://github.com/cvsym/OSMamba. Gehui Li, Bin Chen 0006, Chen Zhao 0002, Lei Zhang 0001, Jian Zhang 0018 |
CVPR | 3 |
| 2025 | SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningabstractMasked video modeling, such as VideoMAE, is an effective paradigm for video self-supervised learning (SSL). However, they are primarily based on reconstructing pixel-level details on natural videos which have substantial temporal redundancy, limiting their capability for semantic representation and sufficient encoding of motion dynamics. To address these issues, this paper introduces a novel SSL approach for video representation learning, dubbed as SMILE, by infusing both spatial and motion semantics. In SMILE, we leverage image-language pretrained models, such as CLIP, to guide the learning process with their high-level spatial semantics. We enhance the representation of motion by introducing synthetic motion patterns in the training data, allowing the model to capture more complex and dynamic content. Furthermore, using SMILE, we establish a new self-supervised video learning paradigm capable of learning strong video representations without requiring any natural video data. We have carried out extensive experiments on 7 datasets with various downstream scenarios. SMILE surpasses current state-of-the-art SSL methods, showcasing its effectiveness in learning more discriminative and generalizable video representations. Code is available: https://github.com/fmthoker/SMILE Fida Mohammad Thoker, Letian Jiang, Chen Zhao 0002, Bernard Ghanem |
CVPR | 3 |
| 2025 | Star: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionabstractImage diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets. Rui Xie 0005, Yinhong Liu, Penghao Zhou, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003, Zhenheng Yang, Ying Tai |
ICCV | 4 |
| 2025 | UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetabstractUltra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle the first challenge, we introduce \textbf{UltraHR-100K}, a high-quality dataset of 100K UHR images with rich captions, offering diverse content and strong visual fidelity. Each image exceeds 3K resolution and is rigorously curated based on detail richness, content complexity, and aesthetic quality. To tackle the second challenge, we propose a frequency-aware post-training method that enhances fine-detail generation in T2I diffusion models. Specifically, we design (i) \textit{Detail-Oriented Timestep Sampling (DOTS)} to focus learning on detail-critical denoising steps, and (ii) \textit{Soft-Weighting Frequency Regularization (SWFR)}, which leverages Discrete Fourier Transform (DFT) to softly constrain frequency components, encouraging high-frequency detail preservation. Extensive experiments on our proposed UltraHR-eval4K benchmarks demonstrate that our approach significantly improves the fine-grained detail quality and overall fidelity of UHR image generation. The code is available at \href{https://github.com/NJU-PCALab/UltraHR-100k}{here}. Chen Zhao 0002, En Ci, Yunzhe Xu, Tiehan Fan, Shanyan Guan, Yanhao Ge, Jian Yang 0003, Ying Tai |
NeurIPS | 1 |
| 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/ Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
Int. J. Comput. Vis. | 80 |
| 2025 | Invertible Diffusion Models for Compressed SensingabstractWhile deep neural networks (NNs) significantly advance image compressed sensing (CS) by improving reconstruction quality, the necessity of training current CS NNs from scratch constrains their effectiveness and hampers rapid deployment. Although recent methods utilize pre-trained diffusion models for image reconstruction, they struggle with slow inference and restricted adaptability to CS. To tackle these challenges, this paper proposes Invertible Diffusion Models (IDM), a novel efficient, end-to-end diffusion-based CS method. IDM repurposes a large-scale diffusion sampling process as a reconstruction model, and fine-tunes it end-to-end to recover original images directly from CS measurements, moving beyond the traditional paradigm of one-step noise estimation learning. To enable such memory-intensive end-to-end fine-tuning, we propose a novel two-level invertible design to transform both 1) multi-step sampling process and 2) noise estimation U-Net in each step into invertible networks. As a result, most intermediate features are cleared during training to reduce up to 93.8% GPU memory. In addition, we develop a set of lightweight modules to inject measurements into noise estimator to further facilitate reconstruction. Experiments demonstrate that IDM outperforms existing state-of-the-art CS networks by up to 2.64 dB in PSNR. Compared to the recent diffusion-based approach DDNM, our IDM achieves up to 10.09 dB PSNR gain and 14.54 times faster inference. Bin Chen 0006, Zhenyu Zhang 0005, Chen Zhao 0002, Jiwen Yu, Shijie Zhao 0001, Jie Chen 0001, Jian Zhang 0018 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 22 |
| 2024 | Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient FinetuningabstractLarge pretrained models are increasingly crucial in modern computer vision tasks. These models are typically used in downstream tasks by end-to-end finetuning, which is highly memory-intensive for tasks with high-resolution data, e.g., video understanding, small object detection, and point cloud analysis. In this paper, we propose Dynamic Reversible Dual-Residual Networks, or Dr2Net, a novel family of network architectures that acts as a surrogate net-work to finetune a pretrained model with substantially re-duced memory consumption. Dr2 Net contains two types of residual connections, one maintaining the residual struc-ture in the pretrained models, and the other making the network reversible. Due to its reversibility, intermediate activations, which can be reconstructed from output, are cleared from memory during training. We use two co-efficients on either type of residual connections respec-tively, and introduce a dynamic training strategy that seam-lessly transitions the pretrained model to a reversible net-work with much higher numerical precision. We evaluate Dr2Net on various pretrained models and various tasks, and show that it can reach comparable performance to con-ventional finetuning but with significantly less memory us-age. Code will be available at https://github.com/coolbay/Dr2Net. Chen Zhao 0002, Shuming Liu 0001, Karttikeya Mangalam, Guocheng Qian, Fatimah Zohra, Abdulmohsen Alghannam, Jitendra Malik, Bernard Ghanem |
CVPR | 1 |
| 2024 | Towards Automated Movie Trailer GenerationabstractMovie trailers are an essential tool for promoting films and attracting audiences. However, the process of creating trailers can be time-consuming and expensive. To stream-line this process, we propose an automatic trailer generation framework that generates plausible trailers from a full movie by automating shot selection and composition. Our approach draws inspiration from machine translation techniques and models the movies and trailers as sequences of shots, thus formulating the trailer generation problem as a sequence-to-sequence task. We introduce Trailer Generation Transformer (TGT), a deep-learning framework utilizing an encoder-decoder architecture. TGT movie encoder is tasked with contextualizing each movie shot representation via self-attention, while the autoregressive trailer de-coder predicts the feature representation of the next trailer shot, accounting for the relevance of shots' temporal order in trailers. Our TGT significantly outperforms previous methods on a comprehensive suite of metrics. Dawit Mureja Argaw, Mattia Soldan, Alejandro Pardo, Chen Zhao 0002, Fabian Caba Heilbron, Joon Son Chung, Bernard Ghanem |
CVPR | 4 |
| 2024 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is un-precedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions-including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
CVPR | 80 |
| 2024 | End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesabstractRecently, temporal action detection (TAD) has seen significant performance improvement with end-to-end training. However, due to the memory bottleneck, only models with limited scales and limited data volumes can afford end-to-end training, which inevitably restricts TAD performance. In this paper, we reduce the memory consumption for end-to-end training, and manage to scale up the TAD backbone to 1 billion parameters and the input video to 1,536 frames, leading to significant detection performance. The key to our approach lies in our proposed temporal-informative adapter (TIA), which is a novel lightweight module that reduces training memory. Using TIA, we free the humongous backbone from learning to adapt to the TAD task by only updating the parameters in TIA. TIA also leads to better TAD representation by temporally aggregating context from adjacent frames throughout the backbone. We evaluate our model across four representative datasets. Owing to our efficient design, we are able to train end-to-end on VideoMAEv2-giant and achieve 75.4% mAP on THUMOS14, being the first end-to-end model to outper-form the best feature-based methods. Code is available at https://github.com/sming256/AdaTAD. Shuming Liu 0001, Chen-Lin Zhang, Chen Zhao 0002, Bernard Ghanem |
CVPR | 3 |
| 2024 | Co2Wounds-V2: Extended Chronic Wounds Dataset from Leprosy PatientsabstractChronic wounds pose an ongoing health concern globally, largely due to the prevalence of conditions such as diabetes and leprosy’s disease. The standard method of monitoring these wounds involves visual inspection by healthcare professionals, a practice that could present challenges for patients in remote areas with inadequate transportation and healthcare infrastructure. This has led to the development of algorithms designed for the analysis and follow-up of wound images, which perform image-processing tasks such as classification, detection, and segmentation. However, the effectiveness of these algorithms heavily depends on the availability of comprehensive and varied wound image data, which is usually scarce. This paper introduces the CO2Wounds-V2 dataset, an extended collection of RGB wound images from leprosy patients with their corresponding semantic segmentation annotations, aiming to enhance the development and testing of image-processing algorithms in the medical field. Karen Sanchez, Carlos Hinojosa, Olinto Mieles, Chen Zhao 0002, Bernard Ghanem, Henry Arguello |
ICIP | 4 |
| 2023 | Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action LocalizationabstractTemporal action localization (TAL) requires long-form reasoning to predict actions of various durations and complex content. Given limited GPU memory, training TAL end to end (i.e., from videos to predictions) on long videos is a significant challenge. Most methods can only train on pre-extracted features without optimizing them for the localization problem, consequently limiting localization performance. In this work, to extend the potential in TAL networks, we propose a novel end-to-end method Re2TAL, which rewires pretrained video backbones for reversible TAL. Re2TAL builds a backbone with reversible modules, where the input can be recovered from the output such that the bulky intermediate activations can be cleared from memory during training. Instead of designing one single type of reversible module, we propose a network rewiring mechanism, to transform any module with a residual connection to a reversible module without changing any parameters. This provides two benefits: (1) a large variety of reversible networks are easily obtained from existing and even future model designs, and (2) the reversible models require much less training effort as they reuse the pre-trained parameters of their original non-reversible versions. Re2TAL, only using the RGB modality, reaches 37.01% average mAP on ActivityNet-v1.3, a new state-of-the-art record, and mAP 64.9% at tIoU=0.5 on THUMOS-14, outperforming all other RGB-only methods. Code is available at https://github.com/coolbay/Re2TAL. Chen Zhao 0002, Shuming Liu 0001, Karttikeya Mangalam, Bernard Ghanem |
CVPR | 1 |
| 2023 | Large-Capacity and Flexible Video Steganography via Invertible Neural NetworkabstractVideo steganography is the art of unobtrusively concealing secret data in a cover video and then recovering the secret data through a decoding protocol at the receiver end. Although several attempts have been made, most of them are limited to low-capacity and fixed steganography. To rectify these weaknesses, we propose a Large-capacity and Flexible Video Steganography Network (LF-VSN) in this paper. For large-capacity, we present a reversible pipeline to perform multiple videos hiding and recovering through a single invertible neural network (INN). Our method can hide/recover 7 secret videos in/from 1 cover video with promising performance. For flexibility, we propose a key-controllable scheme, enabling different receivers to recover particular secret videos from the same cover video through specific keys. Moreover, we further improve the flexibility by proposing a scalable strategy in multiple videos hiding, which can hide variable numbers of secret videos in a cover video with a single model and a single training session. Extensive experiments demonstrate that with the significant improvement of the video steganography performance, our proposed LF-VSN has high security, large hiding capacity, and flexibility. The source code is available at https://github.com/MC-E/LF-VSN. Chong Mou, Youmin Xu, Jiechong Song, Chen Zhao 0002, Bernard Ghanem, Jian Zhang 0018 |
CVPR | 4 |
| 2023 | A Unified Continual Learning Framework with General Parameter-Efficient TuningabstractThe "pre-training → downstream adaptation" presents both new opportunities and challenges for Continual Learning (CL). Although the recent state-of-the-art in CL is achieved through Parameter-Efficient-Tuning (PET) adaptation paradigm, only prompt has been explored, limiting its application to Transformers only. In this paper, we position prompting as one instantiation of PET, and propose a unified CL framework with general PET, dubbed as Learning-Accumulation-Ensemble (LAE). PET, e.g., using Adapter, LoRA, or Prefix, can adapt a pre-trained model to downstream tasks with fewer parameters and resources. Given a PET method, our LAE framework incorporates it for CL with three novel designs. 1) Learning: the pre-trained model adapts to the new task by tuning an online PET module, along with our adaptation speed calibration to align different PET modules, 2) Accumulation: the task-specific knowledge learned by the online PET module is accumulated into an offline PET module through momentum update, 3) Ensemble: During inference, we respectively construct two experts with online/offline PET modules (which are favored by the novel/historical tasks) for prediction ensemble. We show that LAE is compatible with a battery of PET methods and gains strong CL capability. For example, LAE with Adaptor PET surpasses the prior state-of-the-art by 1.3% and 3.6% in last-incremental accuracy on CIFAR100 and ImageNet-R datasets, respectively. Code is available at https://github.com/gqk/LAE. Qiankun Gao, Chen Zhao 0002, Yifan Sun 0003, Teng Xi, Bernard Ghanem, Jian Zhang 0018 |
ICCV | 2 |
| 2023 | EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual QueriesabstractWith the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization (VQ3D). Given an egocentric video clip and an image crop depicting a query object, the goal is to localize the 3D position of the center of that query object with respect to the camera pose of a query frame. Current methods tackle the problem of VQ3D by unprojecting the 2D localization results of the sibling task Visual Queries with 2D Localization (VQ2D) into 3D predictions. Yet, we point out that the low number of camera poses caused by camera re-localization from previous VQ3D methods severally hinders their overall success rate. In this work, we formalize a pipeline (we dub EgoLoc) that better entangles 3D multiview geometry with 2D object retrieval from egocentric videos. Our approach involves estimating more robust camera poses and aggregating multi-view 3D displacements by leveraging the 2D detection confidence, which enhances the success rate of object queries and leads to a significant improvement in the VQ3D baseline performance. Specifically, our approach achieves an overall success rate of up to 87.12%, which sets a new state-of-the-art result in the VQ3D task1. We provide a comprehensive empirical analysis of the VQ3D task and existing solutions, and highlight the remaining challenges in VQ3D. The code is available at https://github.com/Wayne-Mai/EgoLoc. Jinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao 0002, Bernard Ghanem |
ICCV | 4 |
| 2023 | FreeDoM: Training-Free Energy-Guided Conditional Diffusion ModelabstractRecently, conditional diffusion models have gained popularity in numerous applications due to their exceptional generation ability. However, many existing methods are training-required. They need to train a time-dependent classifier or a condition-dependent score estimator, which increases the cost of constructing conditional diffusion models and is inconvenient to transfer across different conditions. Some current works aim to overcome this limitation by proposing training-free solutions, but most can only be applied to a specific category of tasks and not to more general conditions. In this work, we propose a training-Free conditional Diffusion Model (FreeDoM) used for various conditions. Specifically, we leverage off-the-shelf pretrained networks, such as a face detection model, to construct time-independent energy functions, which guide the generation process without requiring training. Furthermore, because the construction of the energy function is very flexible and adaptable to various conditions, our proposed FreeDoM has a broader range of applications than existing training-free methods. FreeDoM is advantageous in its simplicity, effectiveness, and low cost. Experiments demonstrate that FreeDoM is effective for various conditions and suitable for diffusion models of diverse data domains, including image and latent code domains. Code is available at https://github.com/vvictoryuki/FreeDoM. Jiwen Yu, Yinhuai Wang, Chen Zhao 0002, Bernard Ghanem, Jian Zhang 0018 |
ICCV | 3 |
| 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/ Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
CVPR | 20 |
| 2022 | MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsabstractThe recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recent works have begun to discover significant limitations in these datasets, suggesting that state-of-the-art techniques commonly overfit to hidden dataset biases. In this work, we present MAD (Movie Audio Descriptions), a novel benchmark that departs from the paradigm of augmenting existing video datasets with text annotations and focuses on crawling and aligning available audio descriptions of mainstream movies. MAD contains over 384, 000 natural language sentences grounded in over 1, 200 hours of videos and exhibits a significant reduction in the currently diagnosed biases for video-language grounding datasets. MAD's collection strategy enables a novel and more challenging version of video-language grounding, where short temporal moments (typically seconds long) must be accurately grounded in diverse long-form videos that can last up to three hours. We have released MAD's data and baselines code at https://github.com/Soldelli/MAD. Mattia Soldan, Alejandro Pardo, Juan Leon Alcazar, Fabian Caba Heilbron, Chen Zhao 0002, Silvio Giancola, Bernard Ghanem |
CVPR | 5 |
| 2022 | End-to-End Active Speaker Detection
Juan Leon Alcazar, Moritz Cordes, Chen Zhao 0002, Bernard Ghanem |
ECCV (37) | 3 |
| 2022 | R-DFCIL: Relation-Guided Representation Learning for Data-Free Class Incremental Learning
Qiankun Gao, Chen Zhao 0002, Bernard Ghanem, Jian Zhang 0018 |
ECCV (23) | 2 |
| 2021 | Video Self-Stitching Graph Network for Temporal Action LocalizationabstractTemporal action localization (TAL) in videos is a challenging task, especially due to the large variation in action temporal scales. Short actions usually occupy a major proportion in the datasets, but tend to have the lowest performance. In this paper, we confront the challenge of short actions and propose a multi-level cross-scale solution dubbed as video self-stitching graph network (VSGN). We have two key components in VSGN: video self-stitching (VSS) and cross-scale graph pyramid network (xGPN). In VSS, we focus on a short period of a video and magnify it along the temporal dimension to obtain a larger scale. We stitch the original clip and its magnified counterpart in one input sequence to take advantage of the complementary properties of both scales. The xGPN component further exploits the cross-scale correlations by a pyramid of cross-scale graph networks, each containing a hybrid module to aggregate features from across scales as well as within the same scale. Our VSGN not only enhances the feature representations, but also generates more positive anchors for short actions and more short training samples. Experiments demonstrate that VSGN obviously improves the localization performance of short actions as well as achieving the state-of-the-art overall performance on THUMOS-14 and ActivityNet-v1.3. VSGN code is available at https://github.com/coolbay/VSGN. Chen Zhao 0002, Ali K. Thabet, Bernard Ghanem |
ICCV | 1 |
| 2020 | G-TAD: Sub-Graph Localization for Temporal Action DetectionabstractTemporal action detection is a fundamental yet challenging task in video understanding. Video context is a critical cue to effectively detect actions, but current works mainly focus on temporal context, while neglecting semantic context as well as other important context properties. In this work, we propose a graph convolutional network (GCN) model to adaptively incorporate multi-level semantic context into video features and cast temporal action detection as a sub-graph localization problem. Specifically, we formulate video snippets as graph nodes, snippet-snippet correlations as edges, and actions associated with context as target sub-graphs. With graph convolution as the basic operation, we design a GCN block called GCNeXt, which learns the features of each node by aggregating its context and dynamically updates the edges in the graph. To localize each sub-graph, we also design an SGAlign layer to embed each sub-graph into the Euclidean space. Extensive experiments show that G-TAD is capable of finding effective video context without extra supervision and achieves state-of-the-art performance on two detection benchmarks. On ActivityNet-1.3 it obtains an average mAP of 34.09%; on THUMOS14 it reaches 51.6% at [email protected] when combined with a proposal processing method. The code has been made available at https://github.com/frostinassiky/gtad. Mengmeng Xu 0006, Chen Zhao 0002, David S. Rojas, Ali K. Thabet, Bernard Ghanem |
CVPR | 2 |
| 2020 | ThumbNet: One Thumbnail Image Contains All You Need for RecognitionabstractAlthough deep convolutional neural networks (CNNs) have achieved great success in computer vision tasks, its real-world application is still impeded by its voracious demand of computational resources. Current works mostly seek to compress the network by reducing its parameters or parameter-incurred computation, neglecting the influence of the input image on the system complexity. Based on the fact that input images of a CNN contain substantial redundancy, in this paper, we propose a unified framework, dubbed as ThumbNet, to simultaneously accelerate and compress CNN models by enabling them to infer on one thumbnail image. We provide three effective strategies to train ThumbNet. In doing so, ThumbNet learns an inference network that performs equally well on small images as the original-input network on large images. With ThumbNet, not only do we obtain the thumbnail-input inference network that can drastically reduce computation and memory requirements, but also we obtain an image downscaler that can generate thumbnail images for generic classification tasks. Extensive experiments show the effectiveness of ThumbNet, and demonstrate that the thumbnail-input inference network learned by ThumbNet can adequately retain the accuracy of the original-input network even when the input images are downscaled 16 times. Chen Zhao 0002, Bernard Ghanem |
ACM Multimedia | 1 |
| 2018 | BoostNet: A Structured Deep Recursive Network to Boost Image DeblockingabstractTo alleviate the conflict between bit reduction and quality preservation, image deblocking as a post-processing strategy is an attractive and promising solution without changing existing codec. Traditional image deblocking methods mainly rely on manually-crafted image prior models, whereas recent deep network-based methods are usually designed in an inexplicable manner. To combine the merits of both categories of deblocking methods, in this paper, we start from formulating image deblocking as an optimization problem, inspired by which we then design a general and interpretable deep structured network for boosting image deblocking, dubbed BoostNet. Each module of BoostNet correlates to one iteration of the optimization step. In addition, the weights across all modules in BoostNet are shared so that the learnable parameters of BoostNet are tremendously reduced. Furthermore, the quantization matrix, which is not considered in previous network-based methods, is incorporated into our BoostNet as an input. Therefore, one single BoostNet can deal well with various quantization strengths. Experiments demonstrate that our proposed structured deep recursive network BoostNet can produce state-of-the-art deblocking results, while maintaining fast computational speed. Chen Zhao 0002, Jian Zhang 0018, Ronggang Wang, Wen Gao 0001 |
VCIP | 1 |
| 2017 | Video Compressive Sensing Reconstruction via Reweighted Residual SparsityabstractThe compressive sensing (CS) theory indicates that robust reconstruction of signals can be obtained from far fewer measurements than those required by the Nyquist-Shannon theorem. Thus, CS has great potential in video acquisition and processing, considering that it makes the subsequent complex data compression unnecessary. In this paper, we propose a novel algorithm for effectively reconstructing videos from CS measurements. The algorithm comprises double phases, of which the first phase exploits intra-frame correlation and provides good initial recovery for each frame, and the second phase iteratively enhances reconstruction quality by alternating interframe multihypothesis (MH) prediction and sparsity modeling of residuals in a weighted manner. The weights of residual coefficients are updated in each iteration using a statistical method based on the MH predictions. These procedures are performed in the unit of overlapped patches such that potential blocking artifacts can be effectively suppressed through averaging. In addition, we devise an effective scheme based on the split Bregman iteration algorithm to solve the formulated weighted ℓ1minimization problem. The experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods in both objective and subjective reconstruction quality. Chen Zhao 0002, Siwei Ma 0001, Jian Zhang 0018, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Reducing Image Compression Artifacts by Structural Sparse Representation and Quantization Constraint PriorabstractThe block discrete cosine transform (BDCT) has been widely used in current image and video coding standards, owing to its good energy compaction and decorrelation properties. However, because of independent quantization of DCT coefficients in each block, BDCT usually gives rise to visually annoying blocking compression artifacts, especially at low bit rates. In this paper, to reduce blocking artifacts and obtain high-quality images, image deblocking is cast as an optimization problem within maximum a posteriori framework, and a novel algorithm for image deblocking by using structural sparse representation (SSR) prior and quantization constraint (QC) prior is proposed. The SSR prior is utilized to simultaneously enforce the intrinsic local sparsity and the nonlocal self-similarity of natural images, while QC is explicitly incorporated to ensure a more reliable and robust estimation. A new split Bregman iteration-based method with an adaptively adjusted regularization parameter is developed to solve the proposed optimization problem, which makes the entire algorithm more practical. Experiments demonstrate that the proposed image-deblocking algorithm combining SSR and QC outperforms the current state-of-the-art methods in both peak signal-to-noise ratio and visual perception. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Xiaopeng Fan 0001, Yongbing Zhang 0002, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Nonconvex Lp Nuclear Norm based ADMM Framework for Compressed SensingabstractCompressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression methodology. Recent studies further show that image prior models play an important role in image CS recovery. By exploiting the non-local self-similarity of natural images and clustering similar patches, low-rank prior model is adopted in this paper. Different from traditional nuclear norm, we extend thelp(0plpnuclear norm prior model for image CS recovery, which is able to more accurately enforce image structural sparsity and self-similarity at the same time. The proposed optimization problem is efficiently solved within the alternative direction multiplier method (ADMM) framework. Experimental results demonstrate that the proposedlpnuclear norm based ADMM framework for image CS recovery framework exhibits good convergence and achieves significant performance improvements over the current state-of-the-art methods. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |
| 2016 | Compressive-Sensed Image Coding via Stripe-based DPCMabstractThese years have seen the advances of compressive sensing (CS), but efficient coding of sensed measurements is still an issue. In this paper, we propose an image coding system based on the compressive sensing paradigm via stripe-based differential pulse-code modulation (DPCM). In the system, we sample and encode an image in a unit of multiple rows, which we call a stripe. Through extensive experiments, we observe that the correlation between measurements of adjacent stripes are much higher than that of the neighboring blocks. Based on this, we combine the stripe-based CS acquisition with the DPCM framework and design a mechanism that predicts a stripe of measurements from its preceding stripe of measurements. The produced measurement residuals are then quantized and entropy-encoded into binary coding bits, which are tremendously reduced compared to the traditional block-based framework. Furthermore, we provide an image CS reconstruction algorithm corresponding to the stripe-based acquisition. Experiments verify that the reconstruction quality is no worse or even better than the block-based case when much lower bitrate is consumed. In a rate-distortion point of view, the proposed system also outperforms the methods using block-based sampling and achieves the state-of-the-art performance for compressive-sensed image coding. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |
| 2016 | CONCOLOR: Constrained Non-Convex Low-Rank Model for Image DeblockingabstractDue to independent and coarse quantization of transform coefficients in each block, block-based transform coding usually introduces visually annoying blocking artifacts at low bitrates, which greatly prevents further bit reduction. To alleviate the conflict between bit reduction and quality preservation, deblocking as a post-processing strategy is an attractive and promising solution without changing existing codec. In this paper, in order to reduce blocking artifacts and obtain high-quality image, image deblocking is formulated as an optimization problem within maximum a posteriori framework, and a novel algorithm for image deblocking using constrained non-convex low-rank model is proposed. The ℓ(p) (0 < p < 1) penalty function is extended on singular values of a matrix to characterize low-rank prior model rather than the nuclear norm, while the quantization constraint is explicitly transformed into the feasible solution space to constrain the non-convex low-rank optimization. Moreover, a new quantization noise model is developed, and an alternatively minimizing strategy with adaptive parameter adjustment is developed to solve the proposed optimization problem. This parameter-free advantage enables the whole algorithm more attractive and practical. Experiments demonstrate that the proposed image deblocking algorithm outperforms the current state-of-the-art methods in both the objective quality and the perceptual quality. Jian Zhang 0018, Ruiqin Xiong, Chen Zhao 0002, Yongbing Zhang 0002, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Thousand to one: An image compression system via cloud searchabstractWith the advent of the ‘big data’ era, a huge number of images are produced every day. Traditional image compression methods no longer satisfy the demand to store and transmit them. In this paper, we face this challenge and take advantage of the correlations existing between images to achieve a higher compression rate. We propose an image compression system that encodes each image by referencing its correlated images in the cloud. We first extract features from an image and retrieve its similar image from the massive images in the cloud by comparing these features. Then we preprocess the retrieved picture by applying projective transformation and illumination compensation to obtain multiple reference images of higher prediction accuracy. By taking advantage of the redundancy between the reference images and the current image, we encode the current image through prediction coding techniques. The experimental results demonstrate that the proposed method outperforms JPEG and HEVC intra coding by 61.5% and 21.3% on average, respectively. It has an average compression ratio of over a thousand to one. Chen Zhao 0002, Siwei Ma 0001, Wen Gao 0001 |
MMSP | 1 |
| 2015 | A dual structured-sparsity model for compressive-sensed video reconstructionabstractThe compressive sensing theory indicates that robust reconstruction of signals can be obtained from far fewer measurements than those required by the Nyquist theorem. Thus, it has great potential in video acquisition and processing in that it can tremendously save the complex compression required by traditional video coding standards. In this paper, we consider reconstruction of compressive-sensed videos and propose a novel structured-sparsity model with a dual prediction strategy. This structured-sparsity model goes beyond simple sparsity and characterizes the intrinsic structure within the transform coefficients. Also, it exploits the sparsity of the residual between the current patch and its prediction. The prediction process is comprised of a dual strategy, which integrates the advantages of the ambient pixel domain and the measurement domain. In addition, an effective optimization method is designed for solving the formulated problem derived from the model. Experiments demonstrate that the proposed algorithm outperforms the state-of-the art methods for compressive-sensed video reconstruction in both subjective and objective quality. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Ruiqin Xiong, Wen Gao 0001 |
VCIP | 1 |
| 2015 | Adaptive intra-refresh for low-delay error-resilient video coding
Haoming Chen, Chen Zhao 0002, Ming-Ting Sun, Aaron Drake |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Image compressive-sensing recovery using structured laplacian sparsity in DCT domain and multi-hypothesis predictionabstractIn compressive sensing (CS), the seeking of a fair domain is of essentially significance to achieve a high enough degree of signal sparsity. Most methods in the literature, however, use a fixed transform domain or prior information that cannot exhibit enough sparsity for various images. Superiorly, we propose an algorithm to explore the structured Laplacian sparsity of DCT coefficients, which can adapt to the non-stationarity of natural images. Better sparsity is achieved by utilizing the nonlocal similarity of natural images and constructing structured image patch groups. Meanwhile, multiple hypotheses for each pixel could be obtained owing to the overlapping of the structured groups and similar patches. Additionally, for solving the optimization problem formulated from the techniques above, we design an efficient iterative method based on split Bregman iteration (SBI) algorithm. Experimental results demonstrate that the proposed algorithm outperforms the other state-of-the-art methods in both objective and subjective recovery quality. Chen Zhao 0002, Siwei Ma 0001, Wen Gao 0001 |
ICME | 1 |
| 2014 | Video compressive sensing via structured Laplacian modellingabstractSeeking a fair domain in which the signal can exhibit high sparsity is of essential significance in compressive sensing (CS). Most methods in the literature, however, use a fixed transform domain or prior information, which cannot adapt to various video contents. In this paper, we propose a video CS recovery algorithm based on the structured Laplacian model, which can effectually deal with the non-stationarity of natural videos. To build the model, structured patch groups are constructed according to the nonlocal similarity in a temporal scope. By incorporating the model into the CS paradigm, we can formulate an ℓ1-norm optimization problem, for which a solution based on the iterative shrinkage/thresholding algorithms (ISTA) is designed. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods in both objective and subjective recovery quality. Chen Zhao 0002, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 1 |
| 2014 | Image compressive sensing recovery using adaptively learned sparsifying basis via L0 minimization
Jian Zhang 0018, Chen Zhao 0002, Debin Zhao, Wen Gao 0001 |
Signal Process. | 2 |
| 2013 | Wavelet inpainting driven image compression via collaborative sparsity at low bit ratesabstractTo overcome the unsatisfactory encoding quality of conventional image compression methods at low bit rates, the idea of downsampling prior to encoding and upsampling after decoding turns out to be a good solution. Based on this paradigm, we propose a low-bit-rate image compression algorithm by use of the novel wavelet inpainting technique via collaborative sparsity. Superior to the existing methods which operate the sampling in the space domain, we merge the wavelet transform in the downsampling stage, which is verified to be able to preserve much more information. By investigating the local two-dimensional sparsity and the nonlocal three-dimensional sparsity of the image simultaneously, a collaborative sparsity model is exploited to restore the full-resolution image from the decoded downsampled image. Finally a Split Bregman based iterative algorithm is developed to solve the optimization problem. Experimental results demonstrate obvious visual quality improvements, as well as PSNR gains, compared to the state-of-the-art methods under various low bit rates. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 1 |
| 2013 | A highly effective error concealment method for whole frame lossabstractWhen videos are transmitted over the Internet, packet missing is inevitable and an entire frame may get lost. However, most of the literature on error concealment problems can only deal with block loss. For the case of frame loss, they usually fail to achieve satisfactory results. In this paper, we propose a highly effective frame concealment method, which gives amazingly good recovery quality. We base on the theory of multiple hypotheses and devise an adaptive integration scheme to make full use of each hypothesis' strength. Different from the existing methods which mostly rely on motion vectors of previous frames, we fully exploit the correlation between consecutive frames. A novel idea for generating multiple estimates of the lost frame is adopted. Experimental results demonstrate that the proposed algorithm significantly outperforms the state-of-the-art error concealment methods for whole frame loss in both subjective and objective quality. Chen Zhao 0002, Siwei Ma 0001, Jian Zhang 0018, Wen Gao 0001 |
ISCAS | 1 |
| 2012 | Compressed Sensing Recovery via Collaborative SparsityabstractCompressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression approach. Its theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. So one of the most significant challenges in CS is to seek a domain where a signal can exhibit a high degree of sparsity and hence be recovered faithfully. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for compressed sensing recovery via collaborative sparsity (RCoS), which enforces local two-dimensional sparsity and nonlocal three-dimensional sparsity simultaneously in an adaptive hybrid space-transform domain, thus substantially utilizing intrinsic sparsities of natural images and greatly confining the CS solution space. In addition, an efficient augmented Lagrangian based technique is developed to solve the above optimization problem. Experimental results on a wide range of natural images are presented to demonstrate the efficacy of the new CS recovery strategy. Jian Zhang 0018, Debin Zhao, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2012 | Exploiting Image Local and Nonlocal Consistency for Mixed Gaussian-Impulse Noise RemovalabstractMost existing image denoising algorithms can only deal with a single type of noise, which violates the fact that the noisy observed images in practice are often suffered from more than one type of noise during the process of acquisition and transmission. In this paper, we propose a new variational algorithm for mixed Gaussian-impulse noise removal by exploiting image local consistency and nonlocal consistency simultaneously. Specifically, the local consistency is measured by a hyper-Lap lace prior, enforcing the local smoothness of images, while the nonlocal consistency is measured by three-dimensional sparsity of similar blocks, enforcing the nonlocal self-similarity of natural images. Moreover, a Split-Bregman based technique is developed to solve the above optimization problem efficiently. Extensive experiments for mixed Gaussian plus impulse noise show that significant performance improvements over the current state-of-the-art schemes have been achieved, which substantiates the effectiveness of the proposed algorithm. Jian Zhang 0018, Ruiqin Xiong, Chen Zhao 0002, Siwei Ma 0001, Debin Zhao |
ICME | 3 |
| 2012 | Image super-resolution via dual-dictionary learning and sparse representationabstractLearning-based image super-resolution aims to reconstruct high-frequency (HF) details from the prior model trained by a set of high- and low-resolution image patches. In this paper, HF to be estimated is considered as a combination of two components: main high-frequency (MHF) and residual high-frequency (RHF), and we propose a novel image super-resolution method via dual-dictionary learning and sparse representation, which consists of the main dictionary learning and the residual dictionary learning, to recover MHF and RHF respectively. Extensive experimental results on test images validate that by employing the proposed two-layer progressive scheme, more image details can be recovered and much better results can be achieved than the state-of-the-art algorithms in terms of both PSNR and visual perception. Jian Zhang 0018, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Debin Zhao |
ISCAS | 2 |