VLDB 2026 Research / reviewers in the wild / expert
Xiaoyan Sun 0001
dblp:13/1574-1
· DBLP profile ↗
166ranked-venue papers
8as first author
72since 2021 · last 2026
0000-0003-3638-5566ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 146 · 8 first-author · 57 since 2021Artificial intelligence and machine learning · 48 · 35 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing the Unseen: Zooming in the Dark with Event CamerasabstractThis paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these challenges, we present RetinexEVSR, the first event-driven LVSR framework that leverages high-contrast event signals and Retinex-inspired priors to enhance video quality under low-light scenarios. Unlike previous approaches that directly fuse degraded signals, RetinexEVSR introduces a novel bidirectional cross-modal fusion strategy to extract and integrate meaningful cues from noisy event data and degraded RGB frames. Specifically, an illumination-guided event enhancement module is designed to progressively refine event features using illumination maps derived from the Retinex model, thereby suppressing low-light artifacts while preserving high-contrast details. Furthermore, we propose an event-guided reflectance enhancement module that utilizes the enhanced event features to dynamically recover reflectance details via a multi-scale fusion mechanism. Experimental results show that our RetinexEVSR achieves state-of-the-art performance on three datasets. Notably, on the SDSD benchmark, our method can get up to 2.95 dB gain while reducing runtime by 65% compared to prior event-based methods. Dachun Kai, Zeyu Xiao 0002, Huyue Zhu, Jiaxiao Wang, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
AAAI | 6 |
| 2026 | Token-Wise Attention-Guided Semantic Quality Assessment for Compressed Visual Features
Shien Ke, Changsheng Gao, Hadi Amirpour, Zhihua Wang 0002, Xiaoyan Sun 0001 |
QoMEX | 5 |
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 41 |
| 2026 | Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment
Lubin Gan, Jing Zhang 0165, Linhao Qu, Siying Wu, Xiaoyan Sun 0001 |
Expert Syst. Appl. | 6 |
| 2026 | EvTexture++: Event-Driven Texture Enhancement for Video Super-ResolutionabstractEvent-based vision has drawn increasing attention owing to its distinctive properties, including ultra-high temporal resolution and extreme dynamic range. Recent works have introduced it to video super-resolution (VSR) to enhance flow estimation and temporal alignment. In contrast, this paper shifts the focus of event signals from motion refinement to texture enhancement in VSR. We propose EvTexture++, the first event-driven framework dedicated to texture enhancement in VSR. It leverages high-frequency spatiotemporal details from events to improve texture recovery. EvTexture++ incorporates a customized texture enhancement branch, along with an iterative texture enhancement module that progressively exploits high-temporal-resolution event information for texture restoration. This enables gradual refinement of texture regions across iterations, yielding more accurate and detailed high-resolution outputs. Besides intra-frame texture recovery, large motions could degrade inter-frame temporal consistency, particularly in texture regions, leading to texture flickering. To mitigate this, we further exploit the continuous-time motion cues of events to enhance temporal consistency, introducing a temporal texture alignment module that estimates event-guided texture-aware flow for precise inter-frame texture alignment. Moreover, EvTexture++ is designed as a plug-and-play tool to flexibly boost the performance of existing VSR models. Experiments on five datasets demonstrate that EvTexture++ achieves state-of-the-art performance. When integrated into recent VSR models, it yields significant improvements, with gains of up to 1.55 dB in PSNR on the texture-rich Vid4 dataset. Dachun Kai, Jiayao Lu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Panacea+: Panoramic and Controllable Video Generation for Autonomous DrivingabstractThe field of autonomous driving increasingly demands high-quality annotated video training data. In this paper, we propose Panacea+, a powerful and universally applicable framework for generating video data in driving scenes. Built upon the foundation of our previous work, Panacea, Panacea+ adopts a multi-view appearance noise prior mechanism and a super-resolution module for enhanced consistency and increased resolution. Extensive experiments show that the generated video samples from Panacea+ greatly benefit a wide range of tasks on different datasets, including 3D object tracking, 3D object detection, and lane detection tasks on the nuScenes and Argoverse 2 dataset. These results strongly prove Panacea+ to be a valuable data generation framework for autonomous driving. Yuqing Wen, Yingfei Liu, Binyuan Huang, Fan Jia 0006, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Cooperative Multiplex GNN for High-Grade Glioma Survival Prediction From Preoperative Multi-Modal Radiomics-Based Brain NetworksabstractAccurately and preoperatively predicting survival for high-grade gliomas (HGGs) is important for optimizing treatment strategies. Increasing evidence suggests that brain structural and functional connectivity networks derived from advanced magnetic resonance imaging (MRI) are promising predictors for HGG survival. However, advanced MRIs (e.g., diffusion MRI and functional MRI) are generally clinically inaccessible for HGG patients before initiating therapy. To compensate for lack of advanced MRI modalities in brain network studies, in this paper we evaluate the feasibility and performance of predicting HGG survival using exclusively preoperative multi-modal basic structural MRI (sMRI, e.g., T1- and T2-weighted MRI) based brain regional radiomics similarity networks (R2SNs). To this end, we propose a new cooperative multiplex graph neural network (GNN) based multi-modal R2SN integration framework for preoperative HGG survival prediction. First, multi-modal R2SNs are represented by a multiplex network, where each modality-specific R2SN forms one multiplex layer and nodes (i.e., brain regions of interest (ROIs)) are coupled to their replicas across multiplex layers. This facilitates flexible inter-ROI communications both within and between R2SNs. Second, a cooperative GNN is applied to capture intra-modal node feature propagations within each multiplex layer, followed by attention mechanisms used to capture inter-modal node feature interactions across multiplex layers. Finally, a tailored tumor-aware graph pooling is developed to assemble features from the tumor-intersecting ROIs for survival prediction. Extensive experiments on a collected HGG database with three basic sMRI modalities demonstrate the superiority of our method over state-of-the-art baselines in survival stratification. The code is available at https://github.com/ZiLaoTou/TCM-GNN. Ruike Cao, Xingcan Hu, Li Xiao 0002, Gang Qu 0002, Haiye Huo, Vince D. Calhoun, Yu-Ping Wang 0002, Xiaoyan Sun 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Event-Enhanced Blurry Video Super-ResolutionabstractIn this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insufficient motion information for deconvolution and the lack of high-frequency details in LR frames. To address these challenges, we introduce event signals into BVSR and propose a novel event-enhanced network, Ev-DeblurVSR. To effectively fuse information from frames and events for feature deblurring, we introduce a reciprocal feature deblurring module that leverages motion information from intra-frame events to deblur frame features while reciprocally using global scene context from the frames to enhance event features. Furthermore, to enhance temporal consistency, we propose a hybrid deformable alignment module that fully exploits the complementary motion information from inter-frame events and optical flow to improve motion estimation in the deformable alignment process. Extensive evaluations demonstrate that Ev-DeblurVSR establishes a new state-of-the-art performance on both synthetic and real-world datasets. Notably, on real data, our method is 2.59 dB more accurate and 7.28× faster than the recent best BVSR baseline FMA-Net. Dachun Kai, Yueyi Zhang 0001, Jin Wang 0023, Zeyu Xiao 0002, Zhiwei Xiong, Xiaoyan Sun 0001 |
AAAI | 6 |
| 2025 | Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network ApproachabstractEvent cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events, resulting in complex training strategies and increased computational costs. To address these challenges, we propose an efficient hybrid framework for image semantic segmentation, comprising a Spiking Neural Network branch for events and an Artificial Neural Network branch for frames. Specifically, we introduce three specialized modules to facilitate the interaction between these two branches: the Adaptive Temporal Weighting (ATW) Injector, the Event-Driven Sparse (EDS) Injector, and the Channel Selection Fusion (CSF) module. The ATW Injector dynamically integrates temporal features from event data into frame features, enhancing segmentation accuracy by leveraging critical dynamic temporal information. The EDS Injector effectively combines sparse event data with rich frame features, ensuring precise temporal and spatial information alignment. The CSF module selectively merges these features to optimize segmentation performance. Experimental results demonstrate that our framework not only achieves state-of-the-art accuracy across the DDD17-Seg, DSEC-Semantic, and M3ED-Semantic datasets but also significantly reduces energy consumption, achieving a 65% reduction on the DSEC-Semantic dataset. Hebei Li, Yansong Peng, Jiahui Yuan, Peixi Wu, Jin Wang 0023, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
AAAI | 7 |
| 2025 | Spiking Point Transformer for Point Cloud ClassificationabstractSpiking Neural Networks (SNNs) offer an attractive and energy-efficient alternative to conventional Artificial Neural Networks (ANNs) due to their sparse binary activation. When SNN meets Transformer, it shows great potential in 2D image processing. However, their application for 3D point cloud remains underexplored. To this end, we present Spiking Point Transformer (SPT), the first transformer-based SNN framework for point cloud classification. Specifically, we first design Queue-Driven Sampling Direct Encoding for point cloud to reduce computational costs while retaining the most effective support points at each time step. We introduce the Hybrid Dynamics Integrate-and-Fire Neuron (HD-IF), designed to simulate selective neuron activation and reduce over-reliance on specific artificial neurons. SPT attains state-of-the-art results on three benchmark datasets that span both real-world and synthetic datasets in the SNN domain. Meanwhile, the theoretical energy consumption of SPT is at least 6.4x less than its ANN counterpart. Peixi Wu, Bosong Chai, Hebei Li, Menghua Zheng, Yansong Peng, Xuan Nie, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
AAAI | 9 |
| 2025 | Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space ModelabstractBrain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of available modalities for brain tumor segmentation. LS3M excels at efficiently modeling long-range dependencies based on the Mamba design, while incorporating differentiable permutation matrices that reorder input sequences based on modality-specific characteristics. This dynamic reordering ensures that critical spatial inductive biases and long-range semantic correlations inherent in 3D brain MRI are preserved, which is crucial for imcomplete multi-modal brain tumor segmentation. Once the input sequences are reordered using the generated permutation matrix, the Series State Space Model (S3M) block models the relationships between them, capturing both local and long-range dependencies. This enables effective representation of intra-modal and inter-modal relationships, significantly improving segmentation accuracy. Extensive experiments on the BraTS2018 and BraTS2020 datasets demonstrate that LS3M outperforms existing methods, offering a robust solution for brain tumor segmentation, particularly in scenarios with missing modalities. Zheyu Zhang 0002, Yayuan Lu, Feipeng Ma, Yueyi Zhang 0001, Huanjing Yue, Xiaoyan Sun 0001 |
CVPR | 6 |
| 2025 | Entropy-Adapter-Based Deep Image Compression for User-Generated Content with Knowledge DistillationabstractThis study addresses the challenge of domain adaptation in learned image compression, focusing on shifting the model from natural images to user-generated content (UGC) domain. We propose a novel entropy adapter framework augmented with knowledge distillation techniques to improve performance. Unlike existing adapter-based methods that primarily enhance transformation modules, we identify the mismatch between the adapter-based transformation and the fixed entropy network. To resolve this, we introduce adapters within the hypernet and entropy model. Specifically, our decoupled entropy adapter features a deeper residual structure with two independent branches, enabling a separate refinement of mean and scale components. This design improves the accuracy of probability estimation and overall compression efficiency. To further enhance the effectiveness of the adapters, we incorporate a knowledge distillation (KD) strategy with a progressive loss function. It facilitates a smooth transition from KD loss to a rate-distortion (RD) loss in the training process, effectively transferring knowledge from a directly fine-tuned model to the student model. Consequently, this strengthens the adapter's learning capability and improves compression performance. Experimental results show that the proposed method achieves a significant 11.5% bitrate savings compared to the baseline model. Additionally, it demonstrates robust adaptability across diverse network architectures. Yaojun Wu 0001, Chaoyi Lin, Zhipin Deng, Xiaoyan Sun 0001 |
DCC | 5 |
| 2025 | Refining Dataset Distillation via Critical Region Selection and Multiview Teacher GuidanceabstractDataset distillation (DD) aims to improve training efficiency by condensing large datasets into compact yet informative subsets. Existing DD methods primarily use optimization-based approaches for image synthesis, which are computationally intensive. While some recent studies have explored optimization-free alternatives, their simplistic region selection strategies result in poor representations of the original dataset and inefficient utilization of synthetic data during downstream training. To address these limitations, we propose a method called Selection-and-Guidance for Dataset Distillation (SGDD). This approach refines the distillation process through two key stages: region selection and guidance enhancement. Specifically, we first obtain a candidate set of regions from various locations within the images. Then, we utilize the Representative Region Selector and Diverse Region Selector to identify the critical regions for image classification. Furthermore, we generate multiview guidance information for the synthetic data to enhance the distillation process further. By selecting representative and diverse regions while incorporating multiview guidance, our method unleashes the potential of optimization-free DD. Experimental results substantiate the superiority of our approach across various datasets and network architectures. Wenqing Ye, Xiaoyan Sun 0001, Ronald X. Xu, Mingzhai Sun |
ECAI | 5 |
| 2025 | DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene Rendering
Jie Chen 0001, Zhangchi Hu, Peixi Wu, Huyue Zhu, Hebei Li, Xiaoyan Sun 0001 |
ICCV | 6 |
| 2025 | Efficient Spiking Point Mamba for Point Cloud Analysis
Peixi Wu, Bosong Chai, Menghua Zheng, Zhangchi Hu, Jie Chen 0001, Zheyu Zhang 0002, Hebei Li, Xiaoyan Sun 0001 |
ICCV | 9 |
| 2025 | Enhancing Visual Question Answering Via Clustered In-Context Sequence ConfigurationabstractRecent advances in Multimodal In-Context Learning (M-ICL) for Multimodal Large Language Models (MLLMs) have attracted considerable attention. These developments primarily focus on configuring an in-context sequence for a given test case based on instance-level semantic similarity. However, high similarity among demonstrations in the sequence introduces inductive biases, which may mislead MLLMs and ultimately degrade their overall performance. To address this, we propose a novel cluster-based in-context configuration method that adaptively groups candidate data and selects demonstrations from each cluster. This method enhances the diversity within the sequence while preserving semantic consistency, enabling MLLMs to focus on the main intent of the demonstrations. The experimental results on four Visual Question Answering (VQA) benchmarks, including OK-VQA, VQAv2, VizWiz, and TextVQA, demonstrate the effectiveness of our proposed method. Yijun Pan, Hebei Li, Feipeng Ma, Yansong Peng, Siying Wu, Xiaoyan Sun 0001 |
ICIP | 7 |
| 2025 | D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution RefinementabstractWe introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation (GO-LSD). FDR transforms the regression process from predicting fixed coordinates to iteratively refining probability distributions, providing a fine-grained intermediate representation that significantly enhances localization accuracy. GO-LSD is a bidirectional optimization strategy that transfers localization knowledge from refined distributions to shallower layers through self-distillation, while also simplifying the residual prediction tasks for deeper layers. Additionally, D-FINE incorporates lightweight optimizations in computationally intensive modules and operations, achieving a better balance between speed and accuracy. Specifically, D-FINE-L / X achieves 54.0% / 55.8% AP on the COCO dataset at 124 / 78 FPS on an NVIDIA T4 GPU. When pretrained on Objects365, D-FINE-L / X attains 57.1% / 59.3% AP, surpassing all existing real-time detectors. Furthermore, our method significantly enhances the performance of a wide range of DETR models by up to 5.3% AP with negligible extra parameters and training costs. Our code and models: https://github.com/Peterande/D-FINE. Yansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0001 |
ICLR | 5 |
| 2025 | Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple SubjectsabstractDiffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods lack precise layout controllability and overlook the potential of dynamic features of reference subjects in improving fidelity. In this work, we propose Layout-Controllable Personalized Diffusion (LCP-Diffusion) model, a novel framework that integrates subject identity preservation with flexible layout guidance in a tuning-free approach. Our model employs a Dynamic-Static Complementary Visual Refining module to comprehensively capture the intricate details of reference subjects, and introduces a Dual Layout Control mechanism to enforce robust spatial control across both training and inference stages. Extensive experiments validate that LCP-Diffusion excels in both identity preservation and layout controllability. To the best of our knowledge, this is a pioneering work enabling users to "create anything anywhere". Hebei Li, Yansong Peng, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICME | 6 |
| 2025 | DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution TransformationabstractLike image coding in visual data transmission, feature coding is essential for the distributed deployment of large models by significantly reducing transmission and storage burden. However, prior studies have mostly targeted task- or model-specific scenarios, leaving the challenge of universal feature coding across diverse large models largely unexplored. In this paper, we present the first systematic study on universal feature coding for large models. The key challenge lies in the inherently diverse and distributionally incompatible nature of features extracted from different models. For example, features from DINOv2 exhibit highly peaky, concentrated distributions, while those from Stable Diffusion 3 (SD3) are more dispersed and uniform. This distributional heterogeneity severely hampers both compression efficiency and cross-model generalization. To address this, we propose a learned peaky-to-balanced distribution transformation, which reshapes highly skewed feature distributions into a common, balanced target space. This transformation is non-uniform, data-driven, and plug-and-play, enabling effective alignment of heterogeneous distributions without modifying downstream codecs. With this alignment, a universal codec trained on the balanced target distribution can effectively generalize to features from different models and tasks. We validate our approach on three representative large models (LLaMA3, DINOv2, and SD3) across multiple tasks and modalities. Extensive experiments show that our method achieves notable improvements in both compression efficiency and cross-model generalization over task-specific baselines. All source code has been made available at https://github.com/chansongoal/DT-UFC. Changsheng Gao, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001, Weisi Lin |
ACM Multimedia | 5 |
| 2025 | Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object DetectionabstractTiny object detection plays a vital role in drone surveillance, remote sensing, and autonomous systems, enabling the identification of small targets across vast landscapes. However, existing methods suffer from inefficient feature leverage and high computational costs due to redundant feature processing and rigid query allocation. To address these challenges, we propose Dome-DETR, a novel framework with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection. To reduce feature redundancies, we introduce a lightweight Density-Focal Extractor (DeFE) to produce clustered compact foreground masks. Leveraging these masks, we incorporate Masked Window Attention Sparsification (MWAS) to focus computational resources on the most informative regions via sparse attention. Besides, we propose Progressive Adaptive Query Initialization (PAQI), which adaptively modulates query density across spatial areas for better query allocation. Extensive experiments demonstrate that Dome-DETR achieves state-of-the-art performance (+3.3 AP on AI-TOD-V2 and +2.5 AP on VisDrone) while maintaining low computational complexity and a compact model size. Code is available at https://github.com/RicePasteM/Dome-DETR. Zhangchi Hu, Peixi Wu, Jie Chen 0001, Huyue Zhu, Yansong Peng, Hebei Li, Xiaoyan Sun 0001 |
ACM Multimedia | 8 |
| 2025 | EHVC: Efficient Hierarchical Reference and Quality Structure for Neural Video CodingabstractNeural video codecs (NVCs), leveraging the power of end-to-end learning, have demonstrated remarkable coding efficiency improvements over traditional video codecs. Recent research has begun to pay attention to the quality structures in NVCs, optimizing them by introducing explicit hierarchical designs. However, less attention has been paid to the reference structure design, which fundamentally should be aligned with the hierarchical quality structure. In addition, there is still significant room for further optimization of the hierarchical quality structure. To address these challenges in NVCs, we propose EHVC, an efficient hierarchical neural video codec featuring three key innovations: (1) a hierarchical multi-reference scheme that draws on traditional video codec design to align reference and quality structures, thereby addressing the reference-quality mismatch; (2) a lookahead strategy to utilize an encoder-side context from future frames to enhance the quality structure; (3) a layer-wise quality scale with random quality training strategy to stabilize quality structures during inference. With these improvements, EHVC achieves significantly superior performance to the state-of-the-art NVCs. Code will be released in: https://github.com/bytedance/NEVC. Junqi Liao, Yaojun Wu 0001, Chaoyi Lin, Zhipin Deng, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001 |
ACM Multimedia | 7 |
| 2025 | MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language ModelsabstractThe rapid evolution of vision-language models (VLMs), such as CLIP, has demonstrated remarkable zero-shot capabilities in downstream tasks. Prompt learning paradigms like Context Optimization (CoOp) refine learnable prompts for efficient adaptation. However, their application in the biomedical domain remains limited due to the insufficient utilization of specialized biomedical knowledge and cross-modality structural relationships. To address these, we introduce MeDKCoOp, a Medical Dual Knowledge-guided graph adaptation method that leverages systematic integration of knowledge through three aspects: exploit domain-specific knowledge from both textual and visual branches, formalize it into graph-structured representations, and leverage knowledge-guided relation transfer for learning cross-modality fusion. By dynamically optimizing learnable prompts through relation learning process, our method achieves disentangled visual representation and enhances transferability to downstream tasks. Evaluations across 8 biomedical datasets spanning 7 imaging modalities demonstrate state-of-the-art cross-domain generalization, with an average 15.12% accuracy improvement over baselines. Our work establishes a new paradigm through graph prompt learning in medical vision-language models, advancing robust diagnostic AI in data-scarce clinical scenarios. Our code is available at: https://github.com/WangYijun-OUC/MeDKCoOp. Siying Wu, Lubin Gan, Zheyu Zhang 0002, Jing Zhang 0165, Zhangchi Hu, Huyue Zhu, Peixi Wu, Xiaoyan Sun 0001 |
ACM Multimedia | 9 |
| 2025 | MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis
Jing Zhang 0165, Siying Wu, Xun Chen 0001, Yunwei Ou, Xiaoyan Sun 0001 |
Artif. Intell. Medicine | 7 |
| 2025 | Hierarchical Task-aware Temporal Modeling and Matching for few-shot action recognition
Yucheng Zhan, Yijun Pan, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
Neurocomputing | 5 |
| 2025 | Graph relation distillation for efficient biomedical instance segmentation
Xiaoyu Liu 0006, Yueyi Zhang 0001, Zhiwei Xiong, Wei Huang 0036, Bo Hu 0014, Xiaoyan Sun 0001, Feng Wu 0001 |
Pattern Recognit. | 6 |
| 2025 | Semantic-Aware Late-Stage Supervised Contrastive Learning for Fine-Grained Action RecognitionabstractFine-grained action recognition typically faces challenges with lower inter-class variances and higher intra-class variances. Supervised contrastive learning is inherently suitable for this task, as it can decrease intra-class feature distances while increasing inter-class ones. However, directly applying it into fine-grained action recognition encounters two main problems. The first problem stems from the heavy training cost associated with supervised contrastive learning, which requires numerous training epochs, each involving double augmentation views per instance. To address this issue, we propose the late-stage supervised contrastive learning (late-SC) strategy, which effectively reduces the number of training epochs needed for the contrastive learning process. The second problem is that supervised contrastive loss does not explicitly consider the semantic distances between fine-grained actions when adjusting representation distances. This results in less reasonable and efficient adjustments to the representation space. To overcome this limitation, we introduce the semantic-aware temperature adaptation (STA) mechanism, enhancing the suitability of the supervised contrastive loss for fine-grained action recognition. We conduct experiments on several benchmark datasets for fine-grained action recognition, including Epic-Kitchens-55/100, SomethingSomething-V1, and Diving48-V2. The results demonstrate that our proposed method (referred to as LSC-STA) consistently enhances performance across various base feature extractors, without introducing additional inference overhead and incurring only a marginal increase in training expenses. Yijun Pan, Yueyi Zhang 0001, Zilei Wang, Xiaoyan Sun 0001, Feng Wu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | EGVD: Event-Guided Video DerainingabstractRecent research has explored leveraging event cameras, known for their prowess in capturing scenes with nonuniform motion, for video deraining, leading to performance improvements. However, the existing event-based method still faces the challenge that the complex spatiotemporal distribution disrupts temporal information fusion and complicates feature separation. This article proposes a novel end-to-end learning framework for video deraining that effectively extracts the rich dynamic information provided by the event stream. Our framework incorporates two key modules: an event-aware motion detection (EAMD) module that adaptively aggregates multiframe motion information using event-driven masks and a pyramidal adaptive selection module that separates background and rain layers by leveraging contextual priors from both event and conventional camera data. To facilitate efficient training, we introduce a real-world dataset of synchronized rainy videos and event streams. Extensive evaluations on both synthetic and real-world datasets demonstrate the superiority of our proposed method compared to state-of-the-art approaches. The code is available at https://github.com/booker-max/EGVD. Yueyi Zhang 0001, Jin Wang 0023, Wenming Weng, Xiaoyan Sun 0001, Zhiwei Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | IMLS-Splatting: Efficient Mesh Reconstruction from Multi-view Images via Point RepresentationabstractMulti-view mesh reconstruction has long been a challenging problem in graphics and computer vision. In contrast to recent volumetric rendering methods that generate meshes through post-processing, we propose an end-to-end mesh optimization approach called IMLS-Splatting. Our method leverages the sparsity and flexibility of point clouds to efficiently represent the underlying surface. To achieve this, we introduce a splatting-based differentiable Implicit Moving-Least Squares (IMLS) algorithm that enables the fast conversion of point clouds into SDFs and texture fields, optimizing both mesh reconstruction and rasterization. Additionally, the IMLS representation ensures that the reconstructed SDF and mesh maintain continuity and smoothness without the need for extra regularization. With this efficient pipeline, our method enables the reconstruction of highly detailed meshes in approximately 11 minutes, supporting high-quality rendering and achieving state-of-the-art reconstruction performance. Our code is available at https://github.com/SilenKZYoung/IMLS-Splatting. Kaizhi Yang, Liu Dai, Isabella Liu, Xiaoshuai Zhang, Xiaoyan Sun 0001, Xuejin Chen, Zexiang Xu, Hao Su 0001 |
ACM Trans. Graph. | 5 |
| 2024 | TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing ModalitiesabstractNumerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes for missing modalities, inadvertently introducing feature bias and redundant computations. To address these issues, we present the Token Merging transFormer (TMFormer) for robust brain tumor segmentation with missing modalities. TMFormer tackles these challenges by extracting and merging accessible modalities into more compact token sequences. The architecture comprises two core components: the Uni-modal Token Merging Block (UMB) and the Multi-modal Token Merging Block (MMB). The UMB enhances individual modality representation by adaptively consolidating spatially redundant tokens within and outside tumor-related regions, thereby refining token sequences for augmented representational capacity. Meanwhile, the MMB mitigates multi-modal feature fusion bias, exclusively leveraging tokens from present modalities and merging them into a unified multi-modal representation to accommodate varying modality combinations. Extensive experimental results on the BraTS 2018 and 2020 datasets demonstrate the superiority and efficacy of TMFormer compared to state-of-the-art methods when dealing with missing modalities. Zheyu Zhang 0002, Yueyi Zhang 0001, Huanjing Yue, Aiping Liu, Yunwei Ou, Xiaoyan Sun 0001 |
AAAI | 8 |
| 2024 | Image Captioning with Multi-Context Synthetic DataabstractImage captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This potential can be harnessed to create synthetic image-text pairs for training captioning models. Synthetic data can improve cost and time efficiency in data collection, allow for customization to specific domains, bootstrap generalization capability for zero-shot performance, and circumvent privacy concerns associated with real-world data. However, existing methods struggle to attain satisfactory performance solely through synthetic data. We identify the issue as generated images from simple descriptions mostly capture a solitary perspective with limited context, failing to align with the intricate scenes prevalent in real-world imagery. To tackle this, we present an innovative pipeline that introduces multi-context data generation. Beginning with an initial text corpus, our approach employs a large language model to extract multiple sentences portraying the same scene from diverse viewpoints. These sentences are then condensed into a single sentence with multiple contexts. Subsequently, we generate intricate images using the condensed captions through diffusion models. Our model is exclusively trained on synthetic image-text pairs crafted through this process. The effectiveness of our pipeline is validated through experimental results in both the in-domain and cross-domain settings, where it achieves state-of-the-art performance on well-known datasets such as MSCOCO, Flickr30k, and NoCaps. Feipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
AAAI | 5 |
| 2024 | Multi-modal Diffusion Network with Controllable Variability for Medical Image SegmentationabstractIn diffusion-based medical segmentation models, stochastic sampling is commonly used to generate multiple masks. However, the inherent variability in diffusion models can lead to significant biases in some masks, resulting in the fused mask deviating from the true mask. In this study, we propose a novel multi-modal diffusion segmentation network (MMDSN) with controllable variability, specifically designed to address the issue of variability in diffusion models. MMDSN achieves multi-modal conditional control through medical text annotations, thereby enhancing consistency of visual semantic representation and establishing a correspondence between vision and language for diffusion models. Additionally, MMDSN constrains the uncertainty distributions of multiple timesteps within the latent Gaussian space, controlling the variability at each denoising timestep. Extensive experiments on the Qata-Covid19 and MosMed datasets demonstrate that our proposed method surpasses existing state-of-the-art diffusion networks, producing a high-quality, controllable segmentation map with just a single reverse diffusion step and one sampling. Zheyu Zhang 0002, Yueyi Zhang 0001, Jing Zhang 0165, Yunwei Ou, Xiaoyan Sun 0001 |
BIBM | 6 |
| 2024 | Event-Assisted Low-Light Video Object SegmentationabstractIn the realm of video object segmentation (VOS), the challenge of operating under low-light conditions persists, resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras, characterized by their high dynamic range and ability to capture motion information of objects, offer promise in enhancing object visibility and aiding VOS methods under such low-light conditions. This paper introduces a pioneering framework tai-lored for low-light VOS, leveraging event camera data to elevate segmentation accuracy. Our approach hinges on two pivotal components: the Adaptive Cross-Modal Fusion (ACMF) module, aimed at extracting pertinent features while fusing image and event modalities to mitigate noise interference, and the Event-Guided Memory Matching (EGMM) module, designed to rectify the issue of in-accurate matching prevalent in low-light settings. Additionally, we present the creation of a synthetic LLE-DAVIS dataset and the curation of a real-world LLE-vas dataset, encompassing frames and events. Experimental evaluations corroborate the efficacy of our method across both datasets, affirming its effectiveness in low-light scenarios. The datasets are available at https://github.com/HebeiFast/EventLowLightVOS. Hebei Li, Jin Wang 0023, Jiahui Yuan, Wenming Weng, Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001 |
CVPR | 9 |
| 2024 | Scene Adaptive Sparse Transformer for Event-based Object DetectionabstractWhile recent Transformer-based approaches have shown impressive performances on event-based object detection tasks, their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works attempt to reduce these costs by introducing sparse Transformers. However, they display inade-quate sparsity and adaptability when applied to event-based object detection, since these approaches cannot balance the fine granularity of token-level sparsification and the efficiency of window-based Transformers, leading to re-duced performance and efficiency. Furthermore, they lack scene-specific sparsity optimization, resulting in information loss and a lower recall rate. To overcome these limi-tations, we propose the Scene Adaptive Sparse Transformer (SAST). SAST enables window-token co-sparsification, sig-nificantly enhancing fault tolerance and reducing compu-tational overhead. Leveraging the innovative scoring and selection modules, along with the Masked Sparse Window Self-Attention, SAST showcases remarkable scene-aware adaptability: It focuses only on important objects and dy-namically optimizes sparsity level according to scene complexity, maintaining a remarkable balance between performance and computational cost. The evaluation results show that SAST outperforms all other dense and sparse networks in both performance and efficiency on two large-scale event-based object detection datasets (1 Mpx and Genl). Code: https://github.com/Peterande/SAST. Yansong Peng, Hebei Li, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005 |
CVPR | 4 |
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 13 |
| 2024 | Panacea: Panoramic and Controllable Video Generation for Autonomous DrivingabstractThe field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an unlimited numbers of diverse, annotated samples pivotal for autonomous driving advancements. Panacea addresses two critical challenges: ‘Consistency’ and ‘Controllability.’ Consistency ensures temporal and cross-view coherence, while Controllability ensures the alignment of generated content with corresponding annotations. Our approach integrates a novel 4D attention and a two-stage generation pipeline to maintain coherence, supplemented by the ControlNet framework for meticulous control by the Bird'View (BEV) layouts. Extensive qualitative and quantitative evaluations of Panacea on the nuScenes dataset prove its effectiveness in generating high-quality multi-view driving-scene videos. This work notably propels the field of autonomous driving by effectively augmenting the training dataset used for advanced BEV perception techniques. Yuqing Wen, Yingfei Liu, Fan Jia 0006, Chong Luo 0001, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005 |
CVPR | 9 |
| 2024 | Anatomical Consistency Distillation and Inconsistency Synthesis for Brain Tumor Segmentation with Missing ModalitiesabstractMulti-modal Magnetic Resonance Imaging (MRI) is imperative for accurate brain tumor segmentation, offering indispensable complementary information. Nonetheless, the absence of modalities poses significant challenges in achieving precise segmentation. Recognizing the shared anatomical structures between mono-modal and multi-modal representations, it is noteworthy that mono-modal images typically exhibit limited features in specific regions and tissues. In response to this, we present Anatomical Consistency Distillation and Inconsistency Synthesis (ACDIS), a novel framework designed to transfer anatomical structures from multi-modal to mono-modal representations and synthesize modality-specific features. ACDIS consists of two main components: Anatomical Consistency Distillation (ACD) and Modality Feature Synthesis Block (MFSB). ACD incorporates the Anatomical Feature Enhancement Block (AFEB), meticulously mining anatomical information. Simultaneously, Anatomical Consistency ConsTraints (ACCT) are employed to facilitate the consistent knowledge transfer, i.e., the richness of information and the similarity in anatomical structure, ensuring precise alignment of structural features across mono-modality and multi-modality. Complementarily, MFSB produces modality-specific features to rectify anatomical inconsistencies, thereby compensating for missing information in the segmented features. Through validation on the BraTS2018 and BraTS2020 datasets, ACDIS substantiates its efficacy in the segmentation of brain tumors with missing MRI modalities. Zheyu Zhang 0002, Xinzhao Liu, Yueyi Zhang 0001, Huanjing Yue, Yunwei Ou, Xiaoyan Sun 0001 |
ECAI | 7 |
| 2024 | Exploiting Dual-Correlation for Multi-frame Time-of-Flight Denoising
Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ECCV (22) | 3 |
| 2024 | Event-Adapted Video Super-Resolution
Zeyu Xiao 0002, Dachun Kai, Yueyi Zhang 0001, Zhengjun Zha, Xiaoyan Sun 0001, Zhiwei Xiong |
ECCV (42) | 5 |
| 2024 | Event-Based Head Pose Estimation: Benchmark and Method
Jiahui Yuan, Hebei Li, Yansong Peng, Jin Wang 0023, Yuheng Jiang, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ECCV (15) | 7 |
| 2024 | Optimized Decoupled Structure with Non-Local Attention for Deep Image CompressionabstractRecently, a decoupled framework for learning-based image compression has been proposed and adopted into the JPEG AI image coding standard developed by ISO/IEC WG1. The decoupled structure disentangles the sample reconstruction process and the entropy decoding process, making the decoding extremely fast. The corresponding techniques constitute the essential parts of the JPEG AI verification model software. However, its analysis transform and synthesis transform are relatively simple, which are built with stacked convolution layers, thereby may lack the capability to interpret data correlations. In this work, we enhance the transform networks by introducing the non-local attention mechanism, which has proven efficient in image compression tasks. The proposed framework thus shares the merits of the fast decoding from the decoupled architecture and the strong transform capabilities from the non-local attention, making it a stronger candidate for practical end-to-end image codec deployment. Experimental results on the Kodak test set and JPEG AI CfP test set show that our method achieves better BDRate performance compared to the original Decoupled-anchor and significantly faster decoding speed compared to NIC. The proposed solution has been adopted by the IEEE 1857.11 Working Subgroup (1857.11 WSG) in developing neural network-based image coding standards in the 10th Meeting. Xuanye Zhang, Zhaobin Zhang, Yaojun Wu 0001, Semih Esenlik, Xiaoyan Sun 0001, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 5 |
| 2024 | Semantic-Enhanced Point-Box Joint Prompting for Video Object SegmentationabstractThe Segment Anything Model (SAM) has demonstrated outstanding zero-shot performance in image segmentation through efficient point and box prompts. In this paper, we propose a SAM-based Semantic-enhanced Point-Box joint prompting (SAM-SPB) framework for Video Object Segmentation (VOS). SAM-SPB leverages the local structure information and the global semantic cues of interest objects, leading to strong and robust segmentation. To be specific, the local structure information of the objects is maintained by a point tracking branch, and the semantic consistency of the objects across frames are propagated through our proposed semantic-aware memory-based box tracking branch. Compared with previous SAM-based point-centric video segmentation method, we highlight the importance of point-box joint prompting for video object segmentation. The state-of-the-art experimental results on popular VOS benchmarks in the zero-shot setting demonstrate the strong zero-shot ability of the proposed method. Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICIP | 4 |
| 2024 | ESTME: Event-driven Spatio-temporal Motion Enhancement for Micro-Expression RecognitionabstractThe inherently rapid and subtle changes in micro-expressions pose significant challenges for micro-expression recognition (MER). Previous methods, typically relying on frame aggregation or optical flow, struggle to accurately capture subtle changes because of low frame rate. In this paper, we propose an Event-driven Spatio-temporal Motion Enhancement Network, which incorporates event signals captured by an event camera, to assist MER. Specifically, we introduce an Event-Enhanced Motion Extractor module to exploit event signals’ high temporal resolution property, enhancing subtle motion details. We also propose an Event-Guided Attention module to focus on subtle changes in specific areas, capturing more precise spatial features of micro-expressions. Experimental results on synthetic and real-world datasets demonstrate the superiority of our method on MER, showcasing its strong ability to capture subtle motion changes. Peilin Xiao, Yueyi Zhang 0001, Dachun Kai, Yansong Peng, Zheyu Zhang 0002, Xiaoyan Sun 0001 |
ICME | 6 |
| 2024 | EvTexture: Event-driven Texture Enhancement for Video Super-ResolutionabstractEvent-based vision has drawn increasing attention due to its unique characteristics, such as high temporal resolution and high dynamic range. It has been used in video super-resolution (VSR) recently to enhance the flow estimation and temporal alignment. Rather than for motion learning, we propose in this paper the first VSR method that utilizes event signals for texture enhancement. Our method, called EvTexture, leverages high-frequency details of events to better recover texture regions in VSR. In our EvTexture, a new texture enhancement branch is presented. We further introduce an iterative texture enhancement module to progressively explore the high-temporal-resolution event information for texture restoration. This allows for gradual refinement of texture regions across multiple iterations, leading to more accurate and rich high-resolution details. Experimental results show that our EvTexture achieves state-of-the-art performance on four datasets. For the Vid4 dataset with rich textures, our method can get up to 4.67dB gain compared with recent event-based methods. Code: https://github.com/DachunKai/EvTexture. Dachun Kai, Jiayao Lu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICML | 4 |
| 2024 | GRACE: GRadient-based Active Learning with Curriculum Enhancement for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) aims to predict sentiment from text, audio, and visual data of videos. Existing works focus on designing fusion strategies or decoupling mechanisms, which suffer from low data utilization and a heavy reliance on large amounts of labeled data. However, acquiring large-scale annotations for multimodal sentiment analysis is extremely labor-intensive and costly. To address this challenge, we propose GRACE, a GRadient-based Active learning method with Curriculum Enhancement, designed for MSA under a multi-task learning framework. Our approach achieves annotation reduction by strategically selecting valuable samples from the unlabeled data pool while maintaining high-performance levels. Specifically, we introduce informativeness and representativeness criteria, calculated from gradient magnitudes and sample distances, to quantify the active value of unlabeled samples. Additionally, an easiness criterion is incorporated to avoid outliers, considering the relationship between modality consistency and sample difficulty. During the learning process, we dynamically balance sample difficulty and active value, guided by the curriculum learning principle. This strategy prioritizes easier, modality-aligned samples for stable initial training, then gradually increases the difficulty by incorporating more challenging samples with modality conflicts. Extensive experiments demonstrate the effectiveness of our approach on both multimodal sentiment regression and classification benchmarks. Wenqing Ye, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ACM Multimedia | 4 |
| 2024 | Asymmetric Event-Guided Video Super-ResolutionabstractEvent cameras are novel bio-inspired cameras that record asynchronous events with high temporal resolution and dynamic range. Leveraging the auxiliary temporal information recorded by event cameras holds great promise for the task of video super-resolution (VSR). However, existing event-guided VSR methods assume that the event and RGB cameras are strictly calibrated (e.g., pixel-level sensor designs in DAVIS 240/346). This assumption proves limiting in emerging high-resolution devices, such as dual-lens smartphones and unmanned aerial vehicles, where such precise calibration is typically unavailable. To unlock more event-guided application scenarios, we perform the task of asymmetric event-guided VSR for the first time, and we propose an Asymmetric Event-guided VSR Network (AsEVSRN) for this new task. AsEVSRN incorporates two specialized designs for leveraging the asymmetric event stream in VSR. Firstly, the content hallucination module dynamically enhances event and RGB information by exploiting their complementary nature, thereby adaptively boosting representational capacity. Secondly, the event-enhanced bidirectional recurrent cells align and propagate temporal features fused with features from content-hallucinated frames. Within the bidirectional recurrent cells, event-enhanced flow is employed to simultaneously utilize and fuse temporal information at both the feature and pixel levels. Comprehensive experimental results affirm that our method consistently generates superior quantitative and qualitative results. Zeyu Xiao 0002, Dachun Kai, Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ACM Multimedia | 4 |
| 2024 | Visual Perception by Large Language Model's WeightsabstractExisting Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs) and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstrate promising results on various vision-language tasks but are limited by the high computational effort due to the extended input sequence resulting from the involvement of visual tokens. In this paper, instead of input space alignment, we propose a novel parameter space alignment paradigm that represents visual information as model weights. For each input image, we use a vision encoder to extract visual features, convert features into perceptual weights, and merge the perceptual weights with LLM's weights. In this way, the input of LLM does not require visual tokens, which reduces the length of the input sequence and greatly improves efficiency. Following this paradigm, we propose VLoRA with the perceptual weights generator. The perceptual weights generator is designed to convert visual features to perceptual weights with low-rank property, exhibiting a form similar to LoRA. The experimental results show that our VLoRA achieves comparable performance on various benchmarks for MLLMs, while significantly reducing the computational costs for both training and inference. Code and models are released at \url{https://github.com/FeipengMa6/VLoRA}. Feipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang, Fengyun Rao, Shilin Yan, Yueyi Zhang 0001, Siying Wu, Zheng Shou 0001, Xiaoyan Sun 0001 |
NeurIPS | 10 |
| 2024 | Wavelet-like Transform with Subbands Fusion in Decoupled Structure for Deep Image CompressionabstractWavelet-like transform, based on convolutional neural network (CNN), is content-adaptive and has made remarkable achievements in end-to-end image compression. However, the subsequent sequential processing of each subband in the entropy module takes a relatively long decoding time, resulting in incon-venience for real-world applications. In this work, for lossy image compression, the wavelet-like transform is transplanted into the prevailing autoencoder structure to enhance the analysis and synthesis transform due to its excellent decomposition capability. The obtained subbands of different frequencies will undergo a hierarchical decorrelation architecture for subband fusion, also called cross fusing module. The specialized treatment will be applied to different subbands according to their spatial resolution to attain a more compact latent representation. In addition, the proposed solution features an architecture that decouples the arithmetic decoding process from the sample prediction process, which significantly reduces the decoding complexity. Experiments on the Kodak test set show that the proposed method achieves −3.04% BD-Rate compared to existing decoupled end-to-end structure in RGB Peak Signal-to-Noise Ratio (PSNR). Yaojun Wu 0001, Zhaobin Zhang, Semih Esenlik, Xiaoyan Sun 0001, Kai Zhang 0007, Li Zhang 0006 |
PCS | 5 |
| 2024 | Feature Compression With 3D Sparse ConvolutionabstractFeature compression is an important branch of video coding for machines (VCM). While existing methods draw inspiration from image compression, they have not fully utilized the unique characteristics of features. In this paper, we investigate feature characteristics in two key aspects: dimensionality and sparsity. Our analysis reveals that the low spatial dimensionality and high channel dimensionality of features make traditional 2D convolution-based methods, which usually downsample along spatial dimensions while increasing channels, unsuitable for feature compression. To address this, we propose compressing features using 3D convolution. Additionally, considering the sparsity characteristic, we propose applying sparse convolution to reduce model complexity. To thoroughly investigate the proposed 3D sparse convolution-based method, we verify it with various network structures and input features. Experimental results demonstrate the superiority of our proposed method over traditional 2D convolution-based approaches, highlighting its potential for effective feature compression. Changsheng Gao, Qiaoxi Chen, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001 |
VCIP | 6 |
| 2024 | Perceptual Image Compression With Conditional Diffusion TransformersabstractGenerative models have significantly advanced generative AI, particularly in image and video generation. Recognizing their potential, researchers have begun exploring their application in image compression. However, existing methods face two primary challenges: limited performance improvement and high model complexity. In this paper, to address these two challenges, we propose a perceptual image compression solution by introducing a conditional diffusion model. Given that compression performance heavily depends on the decoder’s generative capability, we base our decoder on the diffusion transformer architecture. To address the model complexity problem, we implement the diffusion transformer architecture with Swin transformer. Equipped with enhanced generative capability, we further augment the decoder with informative features using a multi-scale feature fusion module. Experimental results demonstrate that our approach surpasses existing perceptual image compression methods while achieving lower model complexity. Rui Mao 0018, Xinmin Feng, Changsheng Gao, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001 |
VCIP | 6 |
| 2024 | Deep multi-threshold spiking-UNet for image processing
Hebei Li, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001 |
Neurocomputing | 4 |
| 2024 | Multiview hyperedge-aware hypergraph embedding learning for multisite, multiatlas fMRI based functional connectivity network analysis
Wei Wang 0018, Li Xiao 0002, Gang Qu 0002, Vince D. Calhoun, Yu-Ping Wang 0002, Xiaoyan Sun 0001 |
Medical Image Anal. | 6 |
| 2024 | Event-Based Stereo Depth Estimation by Temporal-Spatial Context LearningabstractEvent cameras represent a cutting-edge sensor technology, recording asynchronous pixel-level intensity changes with high temporal resolution and a wide dynamic range. These attributes make event-based stereo depth estimation particularly robust for scenarios characterized by rapid changes and challenging lighting conditions. However, previous learning-based approaches for event-based stereo have often overlooked exploiting the temporal context information within the scene, resulting in suboptimal depth estimations. In this paper, we introduce a novel learning-based network for event-based stereo that incorporates two innovative modules: the Event-based Temporal Aggregation Module (E-TAM) and the Temporal-guided Spatial Context Learning Module (T-SCLM). The E-TAM is designed to capture temporal context information among temporal features extracted from the entire event stream, further the T-SCLM exploits the temporal context information to provide guidance for spatial context learning. Subsequently, these merged features are input into the stereo matching network, ultimately yielding the final disparity map. Experimental evaluations conducted on two real-world datasets affirm the superiority of our method when compared to state-of-the-art approaches. Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Learned Rate-Distortion Cost Prediction for Ultrafast Screen Content Intra CodingabstractAs online collaborations become more prevalent, screen content has become increasingly important in real-time video communications. To reduce communication costs, the H.265/HEVC standard introduced the Screen Content Coding (SCC) extension, which achieves significant bits savings but comes with a higher encoding complexity. There is a need for ultrafast SCC encoding to meet the demands of real-time applications. Our key idea is to predict the rate-distortion (RD) cost of each possible coding unit under each possible mode, rather than performing actual coding to obtain the RD cost. Specifically, we construct neural networks to predict RD costs for intra prediction, palette, and normal intra block copy (IBC) modes. For IBC merge mode, we conduct motion compensation trials and use a linear regression network for prediction. Using the predicted RD costs, we create a partition-mode map set that determines not only block partitioning but also optimal modes, significantly reducing encoding complexity. Our experimental results demonstrate that our method achieves a more than 90% reduction in encoding time with an average 9.4% BD-rate increase compared to the HEVC-SCC reference software in the all-intra configuration. Yanchen Zuo, Changsheng Gao, Dong Liu 0002, Li Li 0040, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Better and Faster: Adaptive Event Conversion for Event-Based Object DetectionabstractEvent cameras are a kind of bio-inspired imaging sensor, which asynchronously collect sparse event streams with many advantages. In this paper, we focus on building better and faster event-based object detectors. To this end, we first propose a computationally efficient event representation Hyper Histogram, which adequately preserves both the polarity and temporal information of events. Then we devise an Adaptive Event Conversion module, which converts events into Hyper Histograms according to event density via an adaptive queue. Moreover, we introduce a novel event-based augmentation method Shadow Mosaic, which significantly improves the event sample diversity and enhances the generalization ability of detection models. We equip our proposed modules on three representative object detection models: YOLOv5, Deformable-DETR, and RetinaNet. Experimental results on three event-based detection datasets (1Mpx, Gen1, and MVSEC-NIGHTL21) demonstrate that our proposed approach outperforms other state-of-the-art methods by a large margin, while achieving a much faster running speed (< 14 ms and < 4 ms for 50 ms event data on the 1Mpx and Gen1 datasets). Yansong Peng, Yueyi Zhang 0001, Peilin Xiao, Xiaoyan Sun 0001, Feng Wu 0001 |
AAAI | 4 |
| 2023 | Paint by Example: Exemplar-based Image Editing with Diffusion ModelsabstractLanguage-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive approach will cause obvious fusing artifacts. We carefully analyze it and propose a content bottleneck and strong augmentations to avoid the trivial solution of directly copying and pasting the exemplar image. Meanwhile, to ensure the controllability of the editing process, we design an arbitrary shape mask for the exemplar image and leverage the classifier-free guidance to increase the similarity to the exemplar image. The whole framework involves a single forward of the diffusion model without any iterative optimization. We demonstrate that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. The code and pretrained models are available at https://github.com/Fantasy-Studio/Paint-by-Example. Binxin Yang, Shuyang Gu, Bo Zhang 0025, Ting Zhang 0002, Xuejin Chen, Xiaoyan Sun 0001, Dong Chen 0003, Fang Wen 0001 |
CVPR | 6 |
| 2023 | GET: Group Event Transformer for Event-Based VisionabstractEvent cameras are a type of novel neuromorphic sensor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity. To address this issue, we propose a novel Group-based vision Transformer backbone for Event-based vision, called Group Event Transformer (GET), which decouples temporal-polarity information from spatial information throughout the feature extraction process. Specifically, we first propose a new event representation for GET, named Group Token, which groups asynchronous events based on their timestamps and polarities. Then, GET applies the Event Dual Self-Attention block, and Group Token Aggregation module to facilitate effective feature communication and integration in both the spatial and temporal-polarity domains. After that, GET can be integrated with different downstream tasks by connecting it with various heads. We evaluate our method on four event-based classification datasets (Cifar10-DVS, N-MNIST, N-CARS, and DVS128Gesture) and two event-based object detection datasets (1Mpx and Gen1), and the results demonstrate that GET outperforms other state-of-the-art methods. The code is available at https://github.com/Peterande/GET-Group-Event-Transformer. Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
ICCV | 4 |
| 2023 | Video Super-Resolution Via Event-Driven Temporal AlignmentabstractVideo super-resolution aims to restore low-resolution videos into their high-resolution counterparts. Existing methods typically rely on optical flow, which assumes linear motion and is sensitive to rapid lighting changes, to capture inter-frame information. Event cameras can asynchronously output high temporal resolution event streams, which can reflect nonlinear motion and are robust to lighting changes. Inspired by these characteristics, we propose an Event-driven Bidirectional Video Super-Resolution (EBVSR) framework. Firstly, we propose an event-assisted temporal alignment module that utilizes events to generate nonlinear motion to align adjacent frames, complementing flow-based methods. Secondly, we build an event-based frame synthesis module that enhances the network’s robustness to lighting changes through a bidirectional cross-modal fusion design. Experimental results on synthetic and real-world datasets demonstrate the superiority of our method. The code is available at https://github.com/DachunKai/EBVSR. Dachun Kai, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICIP | 3 |
| 2023 | Attention-Guided Contrastive Masked Image Modeling for Transformer-Based Self-Supervised LearningabstractSelf-supervised learning with vision transformer (ViT) has gained much attention recently. Most existing methods rely on either contrastive learning or masked image modeling. The former is suitable for global feature extraction but underperforms in fine-grained tasks. The later explores the internal structure of images but ignores the high information sparsity and unbalanced information distribution. In this paper, we propose a new approach called Attention-guided Contrastive Masked Image Modeling (ACoMIM), which integrates the merits of both paradigms and leverages the attention mechanism of ViT for effective representation. Specifically, it has two pretext tasks, predicting the features of masked regions guided by attention and comparing the global features of masked and unmasked images. We show that these two pretext tasks complement each other and improve our method’s performance. The experiments demonstrate that our model transfers well to various downstream tasks such as classification and object detection. Code is available at https://github.com/yczhan/ACoMIM. Yucheng Zhan, Chong Luo 0001, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICIP | 5 |
| 2023 | Multimodal Sentiment Analysis with Preferential Fusion and Distance-aware Contrastive LearningabstractRecent efforts on multimodal sentiment analysis (MSA) leverage data from multiple modalities, among which the text modality is heavily relied on. However, the text modality often contains false correlations between text tokens and sentiment labels, leading to errors in sentiment analysis. To address this issue, we propose a new framework, PriSA, which incorporates the preferential fusion and distance-aware contrastive learning. Specifically, we first propose a preferential inter-modal fusion method, which utilizes the text modality to guide the calculation of the inter-modal correlations. Then the resulting inter-modal features are further used to calculate mixed-modal correlations through our proposed distance-aware contrastive learning, which leverages the distance information of the sentiment labels. At last, we identify the sentiment information based on both the mixed-modal correlations and the discriminative intra-modal features extracted from the visual and audio modalities via a self-attention module. Experimental results show that our proposed PriSA achieves the state-of-the-art performance on four datasets, including MOSEI, MOSI, SIMS, and UR-FUNNY. The code is available at https://github.com/FeipengMa6/PriSA. Feipeng Ma, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICME | 3 |
| 2023 | Overall Survival Time Prediction of Glioblastoma on Preoperative MRI Using Lesion Network Mapping
Xingcan Hu, Li Xiao 0002, Xiaoyan Sun 0001 |
MICCAI (8) | 3 |
| 2023 | EoFormer: Edge-Oriented Transformer for Brain Tumor Segmentation
Dong She, Yueyi Zhang 0001, Zheyu Zhang 0002, Hebei Li, Xiaoyan Sun 0001 |
MICCAI (4) | 6 |
| 2023 | Semantics-Preserving Sketch Embedding for Face GenerationabstractWith recent advances in image-to-image translation tasks, remarkable progress has been witnessed in generating face images from sketches. However, existing methods frequently fail to generate images with details that are semantically and geometrically consistent with the input sketch, especially when various decoration strokes are drawn. To address this issue, we introduce a novel$\mathcal {W}$-$\mathcal {W^+}$encoder architecture to take advantage of the high expressive power of$\mathcal {W^+}$space and semantic controllability of$\mathcal {W}$space. We introduce an explicit intermediate representation for sketch semantic embedding. With a semantic feature matching loss for effective semantic supervision, our sketch embedding precisely conveys the semantics in the input sketches to the synthesized images. Moreover, a novel sketch semantic interpretation approach is designed to automatically extract semantics from vectorized sketches. We conduct extensive experiments on both synthesized sketches and hand-drawn sketches, and the results demonstrate the superiority of our method over existing approaches on both semantics-preserving and generalization ability. Binxin Yang, Xuejin Chen, Chaoqun Wang 0011, Chi Zhang 0044, Xiaoyan Sun 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Exploiting Rigidity Constraints for LiDAR Scene Flow EstimationabstractPrevious LiDAR scene flow estimation methods, especially recurrent neural networks, usually suffer from structure distortion in challenging cases, such as sparse reflection and motion occlusions. In this paper, we propose a novel optimization method based on a recurrent neural network to predict LiDAR scene flow in a weakly supervised manner. Specifically, our neural recurrent network exploits direct rigidity constraints to preserve the geometric structure of the warped source scene during an iterative alignment procedure. An error awarded optimization strategy is proposed to update the LiDAR scene flow by minimizing the point measurement error instead of reconstructing the cost volume multiple times. Trained on two autonomous driving datasets, our network outperforms recent state-of-the-art networks on lidarKITTI by a large margin. The code and models will be available at https://github.com/gtdong-ustc/LiDARSceneFlow. Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
CVPR | 4 |
| 2022 | RPPformer-Flow: Relative Position Guided Point Transformer for Scene Flow EstimationabstractEstimating scene flow for point clouds is one of the key problems in 3D scene understanding and autonomous driving. Recently the point transformer architecture has become a popular and successful solution for 3D computer vision tasks, e.g., point cloud object detection and completion, but its application to scene flow estimation is rarely explored. In this work, we provide a full transformer based solution for scene flow estimation. We first introduce a novel relative position guided point attention mechanism. Then to relax the memory consumption in practice, we provide an efficient implementation of our proposed point attention layer via matrix factorization and nearest neighbor sampling. Finally, we build a pyramid transformer, named RPPformer-Flow, to estimate the scene flow between two consecutive point clouds in a coarse-to-fine manner. We evaluate our RPPformer-Flow on the FlyingThings3D and KITTI Scene Flow 2015 benchmarks. Experimental results show that our method outperforms previous state-of-the-art methods with large margins. Yueyi Zhang 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ACM Multimedia | 4 |
| 2022 | PCGAN: Prediction-Compensation Generative Adversarial Network for MeshesabstractUnlike natural images, the topology similarity among meshes can hardly be handled with classical deep learning because of their irregular structures. Parameterization provides a way to represent meshes in the form of geometry and normal images, which reflects the correlation between neighboring sample locations. Generative Adversarial Networks (GANs) can efficiently generate images without explicitly computing probability densities of the underlying distribution. However, existing GANs such as Coupled Generative Adversarial Network (CoGAN) generally have two drawbacks: (1) Inability to process unnatural images. (2) Insufficient exploration of the inherent relation between normal and the corresponding geometry image. To address these issues, this paper proposes an efficient method named Prediction-Compensation Generative Adversarial Network (PCGAN) to learn a joint distribution of both geometry and normal images, which aims for generating meshes with two GANs. The consistency of two GANs for the geometry and the normal is guaranteed by utilizing a sequence of prediction-compensation pairs. The sequence can estimate the normal image from the geometry image and compensate the geometry from normal progressively. Particularly, the prediction has a closed-form expression, which provides high estimation accuracy and reduces training complexity. Extensive experimental results on facial mesh generation indicate that our PCGAN outperforms CoGAN and other architectures in retaining the geometry of the faces and in generating realistic face meshes with rich facial attributes such as facial expression and morphology. Moreover, quantitative evaluations demonstrate our superior performance compared with the methods mentioned above. Yunhui Shi, Xiaoyan Sun 0001, Jin Wang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Semi-Supervised Neuron Segmentation via Reinforced Consistency LearningabstractEmerging deep learning-based methods have enabled great progress in automatic neuron segmentation from Electron Microscopy (EM) volumes. However, the success of existing methods is heavily reliant upon a large number of annotations that are often expensive and time-consuming to collect due to dense distributions and complex structures of neurons. If the required quantity of manual annotations for learning cannot be reached, these methods turn out to be fragile. To address this issue, in this article, we propose a two-stage, semi-supervised learning method for neuron segmentation to fully extract useful information from unlabeled data. First, we devise a proxy task to enable network pre-training by reconstructing original volumes from their perturbed counterparts. This pre-training strategy implicitly extracts meaningful information on neuron structures from unlabeled data to facilitate the next stage of learning. Second, we regularize the supervised learning process with the pixel-level prediction consistencies between unlabeled samples and their perturbed counterparts. This improves the generalizability of the learned model to adapt diverse data distributions in EM volumes, especially when the number of labels is limited. Extensive experiments on representative EM datasets demonstrate the superior performance of our reinforced consistency learning compared to supervised learning, i.e., up to 400% gain on the VOI metric with only a few available labels. This is on par with a model trained on ten times the amount of labeled data in a supervised manner. Code is available at https://github.com/weih527/SSNS-Net. Wei Huang 0036, Chang Chen 0004, Zhiwei Xiong, Yueyi Zhang 0001, Xuejin Chen, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) targets recognizing new categories by learning transferable image representations. Existing methods find that, by aligning image representations with corresponding semantic labels, the semantic-aligned representations can be transferred to unseen categories. However, supervised by only seen category labels, the learned semantic knowledge is highly task-specific, which makes image representations biased towards seen categories. In this paper, we propose a novel Dual-Contrastive Embedding Network (DCEN) that simultaneously learns task-specific and task-independent knowledge via semantic alignment and instance discrimination. First, DCEN leverages task labels to cluster representations of the same semantic category by cross-modal contrastive learning and exploring semantic-visual complementarity. Besides task-specific knowledge, DCEN then introduces task-independent knowledge by attracting representations of different views of the same image and repelling representations of different images. Compared to high-level seen category supervision, this instance discrimination supervision encourages DCEN to capture low-level visual knowledge, which is less biased toward seen categories and alleviates the representation bias. Consequently, the task-specific and task-independent knowledge jointly make for transferable representations of DCEN, which obtains averaged 4.1% improvement on four public benchmarks. Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Xiaoyan Sun 0001, Houqiang Li |
AAAI | 4 |
| 2021 | Training Spiking Neural Networks with Accumulated Spiking FlowabstractThe fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demonstrate the potential of SNNs in supervised learning by introducing approximated backpropagation (BP) methods. To deal with the non-differentiable spike function in SNNs, these BP methods utilize information from the spatio-temporal domain to adjust the model parameters. With the increasing of time window and network size, the computational complexity of spatio-temporal backpropagation augments dramatically. In this paper, we propose a new backpropagation method for SNNs based on the accumulated spiking flow (ASF), i.e. ASF-BP. In the proposed ASF-BP method, updating parameters does not rely on the spike train of spiking neurons but leverage accumulated inputs and outputs of spiking neurons over the time window, which reduces the BP complexity significantly. We further present an adaptive linear estimation model to approach the dynamic characteristics of spiking neurons statistically. Experimental results demonstrate that with our proposed ASF-BP method, light-weight convolutional SNNs achieve superior performances compared with other spike-based BP methods on both non-neuromorphic (MNIST, CIFAR10) and neuromorphic (CIFAR10-DVS) datasets. The code is available at https://github.com/neural-lab/ASF-BP. Hao Wu 0042, Yueyi Zhang 0001, Wenming Weng, Yongting Zhang, Zhiwei Xiong, Zhengjun Zha, Xiaoyan Sun 0001, Feng Wu 0001 |
AAAI | 7 |
| 2021 | Asymmetric Stereo Color TransferabstractDual-camera systems containing a color camera and a monochrome camera are widely equipped on smartphones. The color camera captures chrominance information while the monochrome camera captures fine details, which causes asymmetry across spectral and spatial dimensions. In these imaging systems, the chrominance information of low-resolution (LR) color images and the spatial information of high-resolution (HR) monochrome images are highly complementary. In this paper, we propose an elaborate convolutional neural network to recover HR color images by transferring color information from LR color images to HR monochrome images. The network contains a novel feature extraction module named U-ASPP and an asymmetric parallax attention module (APAM). Our network achieves state-of-the-art performance on the Flickr1024 stereo dataset with high efficiency. Moreover, the effectiveness of our trained network is validated in real-world asymmetric image pairs captured by a smartphone, which demonstrates that our method has high generalization capability in real-world imaging systems. Jiayong Peng, Yueyi Zhang 0001, Shan Liu 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ICME | 5 |
| 2021 | Uncertainty-Aware Label Rectification for Domain Adaptive Mitochondria Segmentation
Chang Chen 0004, Zhiwei Xiong, Xuejin Chen, Xiaoyan Sun 0001 |
MICCAI (3) | 5 |
| 2021 | Dual Progressive Prototype Network for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) aims to recognize new categories with auxiliary semantic information, e.g., category attributes. In this paper, we handle the critical issue of domain shift problem, i.e., confusion between seen and unseen categories, by progressively improving cross-domain transferability and category discriminability of visual representations. Our approach, named Dual Progressive Prototype Network (DPPN), constructs two types of prototypes that record prototypical visual patterns for attributes and categories, respectively. With attribute prototypes, DPPN alternately searches attribute-related local regions and updates corresponding attribute prototypes to progressively explore accurate attribute-region correspondence. This enables DPPN to produce visual representations with accurate attribute localization ability, which benefits the semantic-visual alignment and representation transferability. Besides, along with progressive attribute localization, DPPN further projects category prototypes into multiple spaces to progressively repel visual representations from different categories, which boosts category discriminability. Both attribute and category prototypes are collaboratively learned in a unified framework, which makes visual representations of DPPN transferable and distinctive.Experiments on four benchmarks prove that DPPN effectively alleviates the domain shift problem in GZSL. Chaoqun Wang 0011, Shaobo Min, Xuejin Chen, Xiaoyan Sun 0001, Houqiang Li |
NeurIPS | 4 |
| 2021 | SMSIR: Spherical Measure Based Spherical Image Representation
Yunhui Shi, Xiaoyan Sun 0001, Jin Wang 0023 |
IEEE Trans. Image Process. | 3 |
| 2020 | Posterior-Guided Neural Architecture SearchabstractThe emergence of neural architecture search (NAS) has greatly advanced the research on network design. Recent proposals such as gradient-based methods or one-shot approaches significantly boost the efficiency of NAS. In this paper, we formulate the NAS problem from a Bayesian perspective. We propose explicitly estimating the joint posterior distribution over pairs of network architecture and weights. Accordingly, a hybrid network representation is presented which enables us to leverage the Variational Dropout so that the approximation of the posterior distribution becomes fully gradient-based and highly efficient. A posterior-guided sampling method is then presented to sample architecture candidates and directly make evaluations. As a Bayesian approach, our posterior-guided NAS (PGNAS) avoids tuning a number of hyper-parameters and enables a very effective architecture sampling in posterior probability space. Interestingly, it also leads to a deeper insight into the weight sharing used in the one-shot NAS and naturally alleviates the mismatch between the sampled architecture and weights caused by the weight sharing. We validate our PGNAS method on the fundamental image classification task. Results on Cifar-10, Cifar-100 and ImageNet show that PGNAS achieves a good trade-off between precision and speed of search among NAS methods. For example, it takes 11 GPU days to search a very competitive architecture with 1.98% and 14.28% test errors on Cifar10 and Cifar100, respectively. Yizhou Zhou, Xiaoyan Sun 0001, Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001 |
AAAI | 2 |
| 2020 | Tracking by Instance Detection: A Meta-Learning ApproachabstractWe consider the tracking problem as a special type of object detection problem, which we call instance detection. With proper initialization, a detector can be quickly converted into a tracker by learning the new instance from a single image. We find that model-agnostic meta-learning (MAML) offers a strategy to initialize the detector that satisfies our needs. We propose a principled three-step approach to build a high-performance tracker. First, pick any modern object detector trained with gradient descent. Second, conduct offline training (or initialization) with MAML. Third, perform domain adaptation using the initial frame. We follow this procedure to build two trackers, named Retina-MAML and FCOS-MAML, based on two modern detectors RetinaNet and FCOS. Evaluations on four benchmarks show that both trackers are competitive against state-of-the-art trackers. On OTB-100, Retina-MAML achieves the highest ever AUC of 0.712. On TrackingNet, FCOS-MAML ranks the first on the leader board with an AUC of 0.757 and the normalized precision of 0.822. Both trackers run in real-time at 40 FPS. Guangting Wang, Chong Luo 0001, Xiaoyan Sun 0001, Zhiwei Xiong, Wenjun Zeng 0001 |
CVPR | 3 |
| 2020 | Spatiotemporal Fusion in 3D CNNs: A Probabilistic ViewabstractDespite the success in still image recognition, deep neural networks for spatiotemporal signal tasks (such as human action recognition in videos) still suffers from low efficacy and inefficiency over the past years. Recently, human experts have put more efforts into analyzing the importance of different components in 3D convolutional neural networks (3D CNNs) to design more powerful spatiotemporal learning backbones. Among many others, spatiotemporal fusion is one of the essentials. It controls how spatial and temporal signals are extracted at each layer during inference. Previous attempts usually start by ad-hoc designs that empirically combine certain convolutions and then draw conclusions based on the performance obtained by training the corresponding networks. These methods only support network-level analysis on limited number of fusion strategies. In this paper, we propose to convert the spatiotemporal fusion strategies into a probability space, which allows us to perform network-level evaluations of various fusion strategies without having to train them separately. Besides, we can also obtain fine-grained numerical information such as layer-level preference on spatiotemporal fusion within the probability space. Our approach greatly boosts the efficiency of analyzing spatiotemporal fusion. Based on the probability space, we further generate new fusion strategies which achieve the state-of-the-art performance on four well-known action recognition datasets. Yizhou Zhou, Xiaoyan Sun 0001, Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001 |
CVPR | 2 |
| 2020 | Enriching Optical Flow with Appearance Information for Action RecognitionabstractOptical flow is a widely used data source for learning motion information, but the complete loss of appearance information limits its ability for action recognition. Therefore we think of enriching optical flow frames with supplementary appearance information to form a new motion data source denoted as Appearance-Supplemented Optical Flow (ASOF). Specifically, we propose a data embedding layer and a stagewise training method to mitigate the scale-wise and density-wise data distribution divergence between the RGB and optical flow frames respectively. We conduct experiments on three benchmark datasets: UCF101 [1], HMDB51 [2] and SomethingSomething-V1 [3]. The results show that our methods can prominently improve optical flow stream recognition accuracy, and further improve the performances of two-stream (score fusion with the RGB stream) with only a little storage increase. Yijun Pan, Xiaoyan Sun 0001, Feng Wu 0001 |
VCIP | 2 |
| 2020 | Learning Redundant Sparsifying Transform based on Equi-Angular FrameabstractDue to the fact that sparse coding in redundant sparse dictionary learning model is NP-hard, interest has turned to the non-redundant sparsifying transform as its sparse coding is computationally cheap. However, natural images typically contain diverse textures that cannot be sparsified well by a non-redundant system. In this paper we propose a new approach for learning redundant sparsifying transform based on equi-angular frame, where the frame and its dual frame are corresponding to applying the forward and the backward transforms. The uniform mutual coherence in the sparsifying transform is enforced by the equi-angular constraint, which better sparsifies diverse textures. In addition, an efficient algorithm is proposed for learning the redundant transform. Experimental results for image representation illustrate the superiority of our proposed method over non-redundant sparsifying transforms. The image denoising results show that our proposed method achieves superior denoising performance, in terms of subjective and objective quality, compared to the K-SVD, the data-driven tight frame method, the learning based sparsifying transform and the overcomplete transform model with block cosparsity (OCTOBOS). Yunhui Shi, Xiaoyan Sun 0001, Nam Ling, Na Qi |
VCIP | 3 |
| 2020 | Temporal-Spatial Mapping for Action RecognitionabstractDeep learning models have enjoyed great success for image related computer vision tasks such as image classification and object detection. For video related tasks such as human action recognition, however, the advancements are not as significant yet. The main challenge is the lack of effective and efficient models in modeling the rich temporal-spatial information in a video. We introduce a simple yet effective operation, termed temporal-spatial mapping, for capturing the temporal evolution of the frames by jointly analyzing all the frames of a video. We propose a video level 2D feature representation by transforming the convolutional features of all frames to a 2D feature map, referred to as VideoMap. With each row being the vectorized feature representation of a frame, the temporal-spatial features are compactly represented, while the temporal dynamic evolution is also well embedded. Based on the VideoMap representation, we further propose a temporal attention model within a shallow convolutional neural network to efficiently exploit the temporal-spatial dynamics. The experiment results show that the proposed scheme achieves state-of-the-art performance, with 4.2% accuracy gain over the temporal segment network, a competing baseline method, on the challenging human action benchmark dataset HMDB51. Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Xiaoyan Sun 0001, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | IENet: Internal and External Patch Matching ConvNet for Web Image Guided DenoisingabstractFrom the non-local self-similarity (NSS)-based image denoising to the convolutional-network (ConvNet)-based image denoising, the denoising performance has been greatly improved. However, it is still not clear how to utilize similar web images to guide image denoising using ConvNet. This paper proposes a novel ConvNet for image denoising to explore both internal (NSS) and external correlations when external similar images are available. Since external similar images may be taken with different viewpoints, focal lengths, and may contain different objects, it is difficult to directly explore external correlations at image level using ConvNet. Therefore, we propose an internal and external patch matching ConvNet (IENet), whose inputs are similar patch cubes extracted from the noisy input and its external similar images. We design three different network structures, namely early-fusion, middle-fusion, and late-fusion of the internal and external cubes to fully combine the strengths of internal and external correlations. The experimental results demonstrate that the proposed method achieves the best denoising results compared with the seven state-of-the-art denoising methods. In specific, the proposed method outperforms the state-of-the-art web image guided denoising method by more than 1 dB on average, which further demonstrates the superiority of the proposed IENet-based filtering over the hand-crafted filtering methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Truong Q. Nguyen, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Image/Video Restoration via Multiplanar Autoregressive Model and Low-Rank OptimizationabstractIn this article, we introduce an image/video restoration approach by utilizing the high-dimensional similarity in images/videos. After grouping similar patches from neighboring frames, we propose to build a multiplanar autoregressive (AR) model to exploit the correlation in cross-dimensional planes of the patch group, which has long been neglected by previous AR models. To further utilize the nonlocal self-similarity in images/videos, a joint multiplanar AR and low-rank based approach is proposed (MARLow) to reconstruct patch groups more effectively. Moreover, for video restoration, the temporal smoothness of the restored video is constrained by the Markov random field (MRF), where MRF encodes a priori knowledge about consistency of patches from neighboring frames. Specifically, we treat different restoration results (from different patch groups) of a certain patch as labels of an MRF, and temporal consistency among these restored patches is imposed. The proposed method is also suitable for other restoration applications such as interpolation and text removal. Extensive experimental results demonstrate that the proposed approach obtains encouraging performance comparing with state-of-the-art methods. Mading Li, Jiaying Liu 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Context-Reinforced Semantic SegmentationabstractRecent efforts have shown the importance of context on deep convolutional neural network based semantic segmentation. Among others, the predicted segmentation map (p-map) itself which encodes rich high-level semantic cues (e.g. objects and layout) can be regarded as a promising source of context. In this paper, we propose a dedicated module, Context Net, to better explore the context information in p-maps. Without introducing any new supervisions, we formulate the context learning problem as a Markov Decision Process and optimize it using reinforcement learning during which the p-map and Context Net are treated as environment and agent, respectively. Through adequate explorations, the Context Net selects the information which has long-term benefit for segmentation inference. By incorporating the Context Net with a baseline segmentation scheme, we then propose a Context-reinforced Semantic Segmentation network (CiSS-Net), which is fully end-to-end trainable. Experimental results show that the learned context brings 3.9% absolute improvement on mIoU over the baseline segmentation method, and the CiSS-Net achieves the state-of-the-art segmentation performance on ADE20K, PASCAL-Context and Cityscapes. Yizhou Zhou, Xiaoyan Sun 0001, Zhengjun Zha, Wenjun Zeng 0001 |
CVPR | 2 |
| 2019 | 3D Mesh Based Inter-Image Prediction for Image Set CompressionabstractA key problem in image set compression is inter-image prediction. Different from the conventional 2D transformation based methods, in this paper we propose a novel 3D mesh based inter-image prediction method. We reconstruct a 3D mesh from the images in the set as a compact representation of the photographed scene. Regarding the images as different projections of the mesh, we build coordinates mappings between images by the multi-view geometry. Exploiting the continuity of the mesh surface, we naturally model the occlusions in the scene and perform inter-image prediction with higher accuracy. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts significantly. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
ICME | 2 |
| 2019 | Quality-Gated Convolutional Lstm for Enhancing Compressed VideoabstractThe past decade has witnessed great success in applying deep learning to enhance the quality of compressed video. However, the existing approaches aim at quality enhancement on a single frame, or only using fixed neighboring frames. Thus they fail to take full advantage of the inter-frame correlation in the video. This paper proposes the Quality-Gated Convolutional Long Short-Term Memory (QG-ConvLSTM) network with bi-directional recurrent structure to fully exploit the advantageous information in a large range of frames. More importantly, due to the obvious quality fluctuation among compressed frames, higher quality frames can provide more useful information for other frames to enhance quality. Therefore, we propose learning the "forget" and "'input" gates in the ConvLSTM cell from quality-related features. As such, the frames with various quality contribute to the memory in ConvLSTM with different importance, making the information of each frame reasonably and adequately used. Finally, the experiments validate the effectiveness of our QG-ConvLSTM approach in advancing the state-of-the-art quality enhancement of compressed video, and the ablation study shows that our QG-ConvLSTM approach is learnt to make a trade-off between quality and correlation when leveraging multi-frame information. The project page: https://github.com/ryangchn/QG-ConvLSTM.git. Xiaoyan Sun 0001, Mai Xu, Wenjun Zeng 0001 |
ICME | 2 |
| 2019 | Mutually Reinforced Spatio-Temporal Convolutional Tube for Human Action RecognitionabstractRecent works use 3D convolutional neural networks to explore spatio-temporal information for human action recognition. However, they either ignore the correlation between spatial and temporal features or suffer from high computational cost by spatio-temporal features extraction. In this work, we propose a novel and efficient Mutually Reinforced Spatio-Temporal Convolutional Tube (MRST) for human action recognition. It decomposes 3D inputs into spatial and temporal representations, mutually enhances both of them by exploiting the interaction of spatial and temporal information and selectively emphasizes informative spatial appearance and temporal motion, meanwhile reducing the complexity of structure. Moreover, we design three types of MRSTs according to the different order of spatial and temporal information enhancement, each of which contains a spatio-temporal decomposition unit, a mutually reinforced unit and a spatio-temporal fusion unit. An end-to-end deep network, MRST-Net, is also proposed based on the MRSTs to better explore spatio-temporal information in human actions. Extensive experiments show MRST-Net yields the best performance, compared to state-of-the-art approaches. Haoze Wu 0003, Jiawei Liu 0001, Zhengjun Zha, Zhenzhong Chen 0001, Xiaoyan Sun 0001 |
IJCAI | 5 |
| 2019 | PGAN: Prediction Generative Adversarial Nets for MeshesabstractUnlike images, the topology similarity among meshes can hardly be handled with traditional signal processing tools because of their irregular structures. Geometry image parameterization provides a way to represent 3D meshes in the form of 2D geometry and normal images. However, most existing methods, including the CoGAN are not suitable for such unnatural images corresponding to meshes. To solve this problem, we propose a Prediction Generative Adversarial Network (PGAN) to learn a joint distribution of geometry and normal images for generating meshes. Particularly, we enforce a prediction constraint on the geometry GAN and normal GAN in our PGAN utilizing the inherent relationship between the geometry and normal. The experimental results on face mesh generation indicate that our PGAN outperforms in generating realistic face models with rich facial attributes such as facial expression and retaining the geometry of the faces. Yunhui Shi, Xiaoyan Sun 0001, Jin Wang 0023 |
VCIP | 3 |
| 2018 | MiCT: Mixed 3D/2D Convolutional Tube for Human Action RecognitionabstractHuman actions in videos are three-dimensional (3D) signals. Recent attempts use 3D convolutional neural networks (CNNs) to explore spatio-temporal information for human action recognition. Though promising, 3D CNNs have not achieved high performance on this task with respect to their well-established two-dimensional (2D) counterparts for visual recognition in still images. We argue that the high training complexity of spatio-temporal fusion and the huge memory cost of 3D convolution hinder current 3D CNNs, which stack 3D convolutions layer by layer, by outputting deeper feature maps that are crucial for high-level tasks. We thus propose a Mixed Convolutional Tube (MiCT) that integrates 2D CNNs with the 3D convolution module to generate deeper and more informative feature maps, while reducing training complexity in each round of spatio-temporal fusion. A new end-to-end trainable deep 3D network, MiCT-Net, is also proposed based on the MiCT to better explore spatio-temporal information in human actions. Evaluations on three well-known benchmark datasets (UCF101, Sport-1M and HMDB-51) show that the proposed MiCT-Net significantly outperforms the original 3D CNNs. Compared with state-of-the-art approaches for action recognition on UCF101 and HMDB51, our MiCT-Net yields the best performance. Yizhou Zhou, Xiaoyan Sun 0001, Zhengjun Zha, Wenjun Zeng 0001 |
CVPR | 2 |
| 2018 | Deep Joint Noise Estimation and Removal for High ISO JPEG ImagesabstractCapturing images under high ISO mode introduces much noise. The statistics of high ISO noise is quite different from that of Gaussian noise. Therefore, this kind of noise is difficult to be removed by traditional Gaussian noise removal methods. This paper proposes a convolutional neural network (CNN) based method to jointly estimate and remove high ISO noise. There are two contributions in this paper. First, we propose a CNN based noise estimation method to estimate the pixel-wise noise level. Due to the Bayer down-sampling process in imaging, the noise variance map is characterized by Bayer patterns. Therefore, we propose packing 2 × 2 blocks in a noisy image into 4D vectors, which makes the pixels with similar noise levels be neighbors. Second, the noise variance map is correlated with the image content. Thus, we propose concatenating the estimated noise variance map with the noisy image, and feed the fused data to the denoising network. The two networks are trained together in an end-to-end fashion. Experimental results demonstrate that the proposed method outperforms state-of-the-art noise estimation and removal methods. Huanjing Yue, Shengdi Zhou, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Chunping Hou |
ICPR | 4 |
| 2018 | Real-Time Object Tracking with Motion InformationabstractMotion is a vital information for object tracking. However, most existing methods, including the classic Siamese FC network [1], only consider the object appearance, and ignore the vital motion feature. In this paper, we design a dual-network object tracker, which is called DOT for short, to effectively combine the appearance and motion information. Our method employs two branches, S-net and M-net, to exploit the appearance and motion information respectively. Moreover, an attention fusion module is also introduced to effectively integrate these two aspects. The experiments carried out on OTB-2013 demonstrate the improvement on object tracking by the integration of motion information with our dual-network and attention fusion. Chaoqun Wang 0011, Xiaoyan Sun 0001, Xuejin Chen, Wenjun Zeng 0001 |
VCIP | 2 |
| 2018 | Multi-Dimensional Sparse ModelsabstractTraditional synthesis/analysis sparse representation models signals in a one dimensional (1D) way, in which a multidimensional (MD) signal is converted into a 1D vector. 1D modeling cannot sufficiently handle MD signals of high dimensionality in limited computational resources and memory usage, as breaking the data structure and inherently ignores the diversity of MD signals (tensors). We utilize the multilinearity of tensors to establish the redundant basis of the space of multi linear maps with the sparsity constraint, and further propose MD synthesis/analysis sparse models to effectively and efficiently represent MD signals in their original form. The dimensional features of MD signals are captured by a series of dictionaries simultaneously and collaboratively. The corresponding dictionary learning algorithms and unified MD signal restoration formulations are proposed. The effectiveness of the proposed models and dictionary learning algorithms is demonstrated through experiments on MD signals denoising, image super-resolution and texture classification. Experiments show that the proposed MD models outperform state-of-the-art 1D models in terms of signal representation quality, computational overhead, and memory storage. Moreover, our proposed MD sparse models generalize the 1D sparse models and are flexible and adaptive to both homogeneous and inhomogeneous properties of MD signals. Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Jingdong Wang 0001, Junbin Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Structure-Revealing Low-Light Image Enhancement Via Robust Retinex ModelabstractLow-light image enhancement methods based on classic Retinex model attempt to manipulate the estimated illumination and to project it back to the corresponding reflectance. However, the model does not consider the noise, which inevitably exists in images captured in low-light conditions. In this paper, we propose the robust Retinex model, which additionally considers a noise map compared with the conventional Retinex model, to improve the performance of enhancing low-light images accompanied by intensive noise. Based on the robust Retinex model, we present an optimization function that includes novel regularization terms for the illumination and reflectance. Specifically, we use norm to constrain the piece-wise smoothness of the illumination, adopt a fidelity term for gradients of the reflectance to reveal the structure details in low-light images, and make the first attempt to estimate a noise map out of the robust Retinex model. To effectively solve the optimization problem, we provide an augmented Lagrange multiplier based alternating direction minimization algorithm without logarithmic transformation. Experimental results demonstrate the effectiveness of the proposed method in low-light image enhancement. In addition, the proposed method can be generalized to handle a series of similar problems, such as the image enhancement for underwater or remote sensing and in hazy or dusty conditions. Mading Li, Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Zongming Guo |
IEEE Trans. Image Process. | 4 |
| 2018 | Photo Stylistic Brush: Robust Style Transfer via Superpixel-Based Bipartite GraphabstractWith the rapid development of social network and multimedia technology, customized image and video stylization have been widely used for various social-media applications. In this paper, we explore the problem of exemplar-based photo style transfer, which provides a flexible and convenient way to invoke fantastic visual impression. Rather than investigating some fixed artistic patterns to represent certain styles as was done in some previous works, our work emphasizes styles related to a series of visual effects in the photograph (e.g., color, tone, and contrast). We propose a photo stylistic brush, an automatic robust style transfer approach based on Super pixel-based BIpartite Graph (SuperBIG). A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixels and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new superpixel-based bipartite graph, and superpixel-level correspondences are generated by bipartite matching. Finally, the refined correspondence guides SuperBIG to perform the transformation in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images. Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Optimal Bit Allocation for CTU Level Rate Control in HEVCabstractFor High Efficiency Video Coding (HEVC), the R–$\lambda $scheme is the latest rate control (RC) scheme, which investigates the relationships among allocated bits, the slope of rate-distortion (R-D) curve$\lambda $, and quantization parameter. However, we argue that bit allocation in the existing R–$\lambda $scheme is not optimal. In this paper, we therefore propose an optimal bit allocation (OBA) scheme for coding tree unit level RC in HEVC. Specifically, to achieve the OBA, we first develop an optimization formulation with a novel R-D estimation, instead of the existing R–$\lambda $estimation. Unfortunately, it is intractable to obtain a closed-form solution to the optimization formulation. We thus propose a recursive Taylor expansion (RTE) method to iteratively solve the formulation. As a result, an approximate closed-form solution can be obtained, thus achieving OBA and bit reallocation. Both theoretical and numerical analyses show the fast convergence speed and little computational time of the proposed RTE method. Therefore, our OBA scheme can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our scheme in three aspects: R-D performance, RC accuracy, and robustness over dynamic scene changes. Shengxi Li, Mai Xu, Zulin Wang, Xiaoyan Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Learning to Detect Video Saliency With HEVC FeaturesabstractSaliency detection has been widely studied to predict human fixations, with various applications in computer vision and image processing. For saliency detection, we argue in this paper that the state-of-the-art High Efficiency Video Coding (HEVC) standard can be used to generate the useful features in compressed domain. Therefore, this paper proposes to learn the video saliency model, with regard to HEVC features. First, we establish an eye tracking database for video saliency detection, which can be downloaded from https://github.com/remega/video_database. Through the statistical analysis on our eye tracking database, we find out that human fixations tend to fall into the regions with large-valued HEVC features on splitting depth, bit allocation, and motion vector (MV). In addition, three observations are obtained with the further analysis on our eye tracking database. Accordingly, several features in HEVC domain are proposed on the basis of splitting depth, bit allocation, and MV. Next, a kind of support vector machine is learned to integrate those HEVC features together, for video saliency detection. Since almost all video data are stored in the compressed form, our method is able to avoid both the computational cost on decoding and the storage cost on raw data. More importantly, experimental results show that the proposed method is superior to other state-of-the-art saliency detection methods, either in compressed or uncompressed domain. Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zhaoting Ye, Zulin Wang |
IEEE Trans. Image Process. | 3 |
| 2017 | Contrast Enhancement Based on Intrinsic Image DecompositionabstractIn this paper, we propose to introduce intrinsic image decomposition priors into decomposition models for contrast enhancement. Since image decomposition is a highly illposed problem, we introduce constraints on both reflectance and illumination layers to yield a highly reliable solution. We regularize the reflectance layer to be piecewise constant by introducing a weighted ℓ1norm constraint on neighboring pixels according to the color similarity, so that the decomposed reflectance would not be affected much by the illumination information. The illumination layer is regularized by a piecewise smoothness constraint. The proposed model is effectively solved by the Split Bregman algorithm. Then, by adjusting the illumination layer, we obtain the enhancement result. To avoid potential color artifacts introduced by illumination adjusting and reduce computing complexity, the proposed decomposition model is performed on the value channel in HSV space. Experiment results demonstrate that the proposed method performs well for a wide variety of images, and achieves better or comparable subjective and objective quality compared with the state-of-the-art methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001, Chunping Hou |
IEEE Trans. Image Process. | 3 |
| 2016 | TenSR: Multi-dimensional Tensor Sparse RepresentationabstractThe conventional sparse model relies on data representation in the form of vectors. It represents the vector-valued or vectorized one dimensional (1D) version of an signal as a highly sparse linear combination of basis atoms from a large dictionary. The 1D modeling, though simple, ignores the inherent structure and breaks the local correlation inside multidimensional (MD) signals. It also dramatically increases the demand of memory as well as computational resources especially when dealing with high dimensional signals. In this paper, we propose a new sparse model TenSR based on tensor for MD data representation along with the corresponding MD sparse coding and MD dictionary learning algorithms. The proposed TenSR model is able to well approximate the structure in each mode inherent in MD signals with a series of adaptive separable structure dictionaries via dictionary learning. The proposed MD sparse coding algorithm by proximal method further reduces the computational cost significantly. Experimental results with real world MD signals, i.e. 3D Multi-spectral images, show the proposed TenSR greatly reduces both the computational and memory costs with competitive performance in comparison with the state-of-the-art sparse representation methods. We believe our proposed TenSR model is a promising way to empower the sparse representation especially for large scale high order signals. Na Qi, Yunhui Shi, Xiaoyan Sun 0001 |
CVPR | 3 |
| 2016 | MARLow: A Joint Multiplanar Autoregressive and Low-Rank Approach for Image Completion
Mading Li, Jiaying Liu 0001, Zhiwei Xiong, Xiaoyan Sun 0001, Zongming Guo |
ECCV (7) | 4 |
| 2016 | Internal-video mode dependent directional transformabstractAs the projection of the real world, videos usually have many repeated patterns with similar structures cross regions, presenting strong non-local correlations. Moreover, different videos own different characteristics. Exploitation of the non-local correlations by off-line training of transforms has attracted considerable attention over the past years for compression. However, the samples used for training the transforms are usually collected from a predefined set of training videos to avoid the transmission of transform matrixes. There is no guarantee that the characteristics of those training videos and the corresponding transforms fit the current coding video very well. To address that, this paper proposes an internal-video mode dependent directional transform in HEVC for intra coding. In this scheme, for coding a video clip, based on different directions, a set of sample blocks is collected for each direction to train a Karhunen-Loeve transform at the video clip level. During encoding, the better transform between the proposed transform and original DCT/DST in HEVC is determined based on rate-distortion optimization. These transform matrixes are entropy coded to the bitstream. The proposed method can capture the statistical characteristics of the coded video clip and provide efficient transforms. Experimental results show that the proposed method achieves significant performance improvements in comparison with HEVC for intra coding, up to 13% BD-rate savings for the all intra configuration. Cuiling Lan, Yunhui Shi, Wenpeng Ding, Xiaoyan Sun 0001 |
VCIP | 5 |
| 2016 | Subjective-Driven Complexity Control Approach for HEVCabstractThe latest High Efficiency Video Coding (HEVC) standard significantly increases the encoding complexity for improving its coding efficiency, compared with the preceding H.264/Advanced Video Coding (AVC) standard. In this paper, we present a novel subjective-driven complexity control (SCC) approach to reduce and control the encoding complexity of HEVC. Through reasonably adjusting the maximum depth of each largest coding unit (LCU), the encoding complexity can be reduced to a target level with minimal visual distortion. Specifically, the maximum depths of different LCUs can be varied through solving the proposed optimization formulation of complexity control, based on two explored relationships: 1) the relationship between the maximum depth and encoding complexity and 2) the relationship between the maximum depth and visual distortion. Besides, the subjective visual quality is favored with a novel subjective-driven constraint imposed in the formulation, on the basis of a visual attention model. Finally, the experimental results show that our approach can achieve a wide range of encoding complexity control (as low as 20%) for HEVC, with the smallest complexity bias being 0.2%. Meanwhile, our SCC approach outperforms other two state-of-the-art complexity control approaches, in terms of both control accuracy and visual quality. Xin Deng 0002, Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zulin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Lossless Compression of JPEG Coded Photo CollectionsabstractThe explosion of digital photos has posed a significant challenge to photo storage and transmission for both personal devices and cloud platforms. In this paper, we propose a novel lossless compression method to further reduce the size of a set of JPEG coded correlated images without any loss of information. The proposed method jointly removes inter/intra image redundancy in the feature, spatial, and frequency domains. For each collection, we first organize the images into a pseudo video by minimizing the global prediction cost in the feature domain. We then present a hybrid disparity compensation method to better exploit both the global and local correlations among the images in the spatial domain. Furthermore, the redundancy between each compensated signal and the corresponding target image is adaptively reduced in the frequency domain. Experimental results demonstrate the effectiveness of the proposed lossless compression method. Compared with the JPEG coded image collections, our method achieves average bit savings of more than 31%. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | DAC-Mobi: Data-Assisted Communications of Mobile Images with Cloud Computing SupportabstractThis research proposes a novel data assisted image transmission scheme, which utilizes a large amount of correlated images stored in the cloud to improve the spectrum efficiency and visual quality. First, a two-layer Coset coding is proposed for the DCT coefficients transmission. The most significant bits (MSB) of the coefficients are generated by the first layer Coset and together with a few low frequency coefficients are transmitted through the most reliable channel coding and digital modulation. The middle bits generated by the second layer Coset are discarded by the sender and the residual bits are transmitted through amplitude modulation. Based on the MSB and the residual bits, an approximation of the original image is reconstructed. With this approximation, a lot of correlated images can be retrieved from the cloud, which are used to recover the discarded middle bits. The two layer Coset coding can significantly decrease the data energy so as to improve the transmission power efficiency. Hence, the end to end distortion of amplitude modulation can be reduced. Second, the image quality can be further improved by joint internal and external denoising with the retrieved images. Simulations show that the proposed scheme outperforms conventional digital schemes about 4 dB in peak signal to noise power ratio (PSNR) and achieves 2 dB gain over the state-of-the-art uncoded transmission. At low signal to noise power ratio (SNR), an additional 2-3 dB gain is achieved. The visual quality comparison also validates the objective image assessment result. Jun Wu 0006, Jian Wu 0022, Hao Cui 0001, Chong Luo 0001, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2015 | Single image super-resolution via 2D sparse representationabstractImage super-resolution with sparsity prior provides promising performance. However, traditional sparse-based super resolution methods transform a two dimensional (2D) image into a one dimensional (1D) vector, which ignores the intrinsic 2D structure as well as spatial correlation inherent in images. In this paper, we propose the first image super-resolution method which reconstructs a high resolution image from its low resolution counterpart via a two dimensional sparse model. Correspondingly, we present a new dictionary learning algorithm to fully make use of the corresponding relationship of two pairs of 2D dictionaries of low and high resolution images, respectively. Experimental results demonstrate that our proposed image super-resolution with 2D sparse model outperforms state-of-the-art 1D sparse model based super resolution methods in terms of both reconstruction ability and memory usage. Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Wenpeng Ding |
ICME | 3 |
| 2015 | Enhancing nighttime surveillance video via gradient fusionabstractThis paper presents an effective method to enhance the quality of dim light surveillance via gradient fusion. We simply take the advantage that surveillance cameras capture a large quantity of valuable information at the same viewpoint during the day. And it can be used to make the video at night easier to perceive. Based on a gradient domain technique, all the important local perceptual cues from the original video are automatically combined with the supporting daytime context, while avoiding traditional problems such as aliasing, ghosting and haloing. Experimental results show that our method outperforms the state-of-the-art ones, and can even handle some challenging conditions without altering the parameters. Xiaoyan Sun 0001, Feng Wu 0001 |
VCIP | 2 |
| 2015 | Single image super-resolution via 2D nonlocal sparse representationabstractImage super-resolution based on sparse model with patch clustering and nonlocal similarity provides promising performance. However, the traditional one dimensional (1D) sparse model enforces a 1D dictionary for every cluster of patches to capture complex structures and different features in images. The total dictionary will take expensive memory, which can be alleviated at cost of representation power. Recently, two dimensional (2D) sparse model has been proved to efficiently represent images and save memory usage. In this paper, we propose to integrate 2D sparse model with patch clustering and nonlocal similarity into a variational framework as 2D nonlocal sparse representation (2DNSR) for image SR to save memory cost and ensure SR performances. We also present a 2DNSR algorithm for image SR where each group of similar patches decompose on the respective 2D dictionaries. Experimental results on image SR demonstrate our proposed 2D nonlocal representation outperforms 2D sparse model and achieves competitive performance to state-of-the-art 1D nonlocal sparse models whereas with much less memory costs. Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Wenpeng Ding |
VCIP | 3 |
| 2015 | 2D nonlocal sparse representation for image denoisingabstractTwo dimensional (2D) sparse representation provides promising performance in image denoising by cooperatively exploiting horizontal and vertical features inherent in images by two dictionaries. In this paper, we first propose integrating the 2D sparse model with clustering and nonlocal regularization into a unified variational framework, defined as 2D nonlocal sparse representation (2DNSR), for optimization. Within this framework, we then present a dictionary learning method for image denoising which jointly decomposes groups of similar noisy patches on subsets of 2D dictionaries. We finally present a 2DNSR-based algorithm for image denoising. Experimental results on image denoising show our proposed 2D nonlocal sparse representation outperforms the 2D sparse model and achieves competitive performance to state-of-the-art nonlocal sparse models whereas with much less memory costs. Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Wenpeng Ding |
VCIP | 3 |
| 2015 | Incremental SfM based lossless compression of JPEG coded photo albumabstractThe key problem in photo album compression is how to exploit the correlation among the images. In this paper, we propose a novel incremental structure from motion (SfM) based prediction method for lossless photo album compression. Unlike the previous methods, we exploit the redundancy among images through their inherent geometric relationship generated by SfM. Based on the point cloud and camera poses, each prediction image is generated by projecting, triangulation and warping. Finally, the target image is compressed by an HEVC-like encoder with the prediction image as main reference. Experimental results demonstrate the advantage of our method, especially for images with scenes containing complicated geometric structures. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
VCIP | 2 |
| 2015 | Deblurring Saturated Night Image With Function-Form KernelabstractDeblurring saturated night images are a challenging problem because such images have low contrast combined with heavy noise and saturated regions. Unlike the deblurring schemes that discard saturated regions when estimating blur kernels, this paper proposes a novel scheme to deduce blur kernels from saturated regions via a novel kernel representation and advanced algorithms. Our key technical contribution is the proposed function-form representation of blur kernels, which regularizes existing matrix-form kernels using three functional components: 1) trajectory; 2) intensity; and 3) expansion. From automatically detected saturated regions, their skeleton, brightness, and width are fitted into the corresponding three functional components of blur kernels. Such regularization significantly improves the quality of kernels deduced from saturated regions. Second, we propose an energy minimizing algorithm to select and assign the deduced function-form kernels to partitioned image regions as the initialization for non-uniform deblurring. Finally, we convert the assigned function-form kernels into matrix form for more detailed estimation in a multi-scale deconvolution. Experimental results show that our scheme outperforms existing schemes on challenging real examples. Xiaoyan Sun 0001, Lu Fang 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Image Denoising by Exploring External and Internal CorrelationsabstractSingle image denoising suffers from limited data collection within a noisy image. In this paper, we propose a novel image denoising scheme, which explores both internal and external correlations with the help of web images. For each noisy patch, we build internal and external data cubes by finding similar patches from the noisy and web images, respectively. We then propose reducing noise by a two-stage strategy using different filtering approaches. In the first stage, since the noisy patch may lead to inaccurate patch selection, we propose a graph based optimization method to improve patch matching accuracy in external denoising. The internal denoising is frequency truncation on internal cubes. By combining the internal and external denoising patches, we obtain a preliminary denoising result. In the second stage, we propose reducing noise by filtering of external and internal cubes, respectively, on transform domain. In this stage, the preliminary denoising result not only enhances the patch matching accuracy but also provides reliable estimates of filtering parameters. The final denoising image is obtained by fusing the external and internal filtering results. Experimental results show that our method constantly outperforms state-of-the-art denoising schemes in both subjective and objective quality measurements, e.g., it achieves >2 dB gain compared with BM3D at a wide range of noise levels. Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Separable Kernel for Image DeblurringabstractIn this paper, we deal with the image deblurring problem in a completely new perspective by proposing separable kernel to represent the inherent properties of the camera and scene system. Specifically, we decompose a blur kernel into three individual descriptors (trajectory, intensity and point spread function) so that they can be optimized separately. To demonstrate the advantages, we extract one-pixel-width trajectories of blur kernels and propose a random perturbation algorithm to optimize them but still keeping their continuity. For many cases, where current deblurring approaches fall into local minimum, excellent deblurred results and correct blur kernels can be obtained by individually optimizing the kernel trajectories. Our work strongly suggests that more constraints and priors should be introduced to blur kernels in solving the deblurring problem because blur kernels have lower dimensions than images. Lu Fang 0001, Feng Wu 0001, Xiaoyan Sun 0001, Houqiang Li |
CVPR | 4 |
| 2014 | CID: Combined Image Denoising in Spatial and Frequency Domains Using Web ImagesabstractIn this paper, we propose a novel two-step scheme to filter heavy noise from images with the assistance of retrieved Web images. There are two key technical contributions in our scheme. First, for every noisy image block, we build two three dimensional (3D) data cubes by using similar blocks in retrieved Web images and similar nonlocal blocks within the noisy image, respectively. To better use their correlations, we propose different denoising strategies. The denoising in the 3D cube built upon the retrieved images is performed as median filtering in the spatial domain, whereas the denoising in the other 3D cube is performed in the frequency domain. These two denoising results are then combined in the frequency domain to produce a denoising image. Second, to handle heavy noise, we further propose using the denoising image to improve image registration of the retrieved Web images, 3D cube building, and the estimation of filtering parameters in the frequency domain. Afterwards, the proposed denoising is performed on the noisy image again to generate the final denoising result. Our experimental results show that when the noise is high, the proposed scheme is better than BM3D by more than 2 dB in PSNR and the visual quality improvement is clear to see. Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
CVPR | 2 |
| 2014 | Lossless compression of JPEG coded photo albumsabstractThe explosion in digital photography poses a significant challenge when it comes to photo storage for both personal devices and the Internet. In this paper, we propose a novel lossless compression method to further reduce the storage size of a set of JPEG coded correlated images. In this method, we propose jointly removing the inter-image redundancy in the feature, spatial, and frequency domains. For each album, we first organize the images into a pseudo video by minimizing the global predictive cost in the feature domain. We then introduce a disparity compensation method to enhance the spatial correlation between images. Finally, the redundancy between the compensated signal and the corresponding target image is adaptively reduced in the frequency domain. Moreover, our proposed scheme is able to losslessly recover not only raw images but also JPEG files. Experimental results demonstrate the efficiency of our proposed lossless compression, which achieves more than 12% bit-saving on average compared with JPEG coded albums. Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
VCIP | 2 |
| 2014 | Learning adaptive filter banks for hierarchical image representationabstractConventional hierarchical image representation methods, e.g. Wavelet transform, use pre-determined filter banks which lack in adaption to the variant statistical characteristics of images. In this paper, we propose learning adaptive filter banks for hierarchical sparse image representation with a wavelet-like compact form using a deconvolutional network. The proposed scheme is verified by evaluating its sparsity in image representation. Experimental results demonstrate that the proposed scheme outperforms 9/7 and 5/3 wavelets transform in terms of both objective and subjective qualities under the same sparsity. Yunhui Shi, Wenpeng Ding, Xiaoyan Sun 0001 |
VCIP | 4 |
| 2013 | Large scale image retrieval with visual groupsabstractBag-of-visual words (BoW) representation has been widely used in the large scale image retrieval. Though efficient, it ignores the geometric correlation among visual words, whereas the geometric verification has demonstrated its effectiveness in image retrieval. In this paper, we propose a new representation - visual group, to improve the retrieval precision by grouping the geometrically related features based on the inclusion relationship between features at different scales. A visual group consists of a master feature and several member features covered by the master feature. The geometric constraint inside each group is introduced into visual group matching for efficient geometric verification. Experimental evaluation on the dataset Oxford5K+Flickr1M shows that our visual group based image search approach outperforms BoW and the state-of-the-art visual phrase based schemes. Lican Dai, Xiaoyan Sun 0001, Feng Wu 0001, Nenghai Yu |
ICIP | 2 |
| 2013 | Two dimensional analysis sparse modelabstractAn analysis sparse model represents an image signal by multiplying it using an analysis dictionary, leading to a sparse outcome. It transforms an image (two dimensional signal) into a one-dimensional (1D) vector. However, this 1D model ignores the two dimensional property and breaks the local spatial correlation inside images. In this paper, we propose a two dimensional (2D) analysis sparse model. Our 2D model uses two analysis dictionaries to efficiently exploit the horizontal and vertical features simultaneously. The corresponding sparse coding and dictionary learning algorithm are also presented in this paper. The 2D sparse model is further evaluated for image denoising. Experimental results demonstrate our 2D analysis sparse model outperforms a state-of-the-art 1D analysis model in terms of both denoising ability and memory usage. Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Jingdong Wang 0001, Wenpeng Ding |
ICIP | 3 |
| 2013 | Two dimensional synthesis sparse modelabstractSparse representation has been proved to be very efficient in machine learning and image processing. Traditional image sparse representation formulates an image into a one dimensional (1D) vector which is then represented by a sparse linear combination of the basis atoms from a dictionary. This 1D representation ignores the local spatial correlation inside one image. In this paper, we propose a two dimensional (2D) sparse model to much efficiently exploit the horizontal and vertical features which are represented by two dictionaries simultaneously. The corresponding sparse coding and dictionary learning algorithm are also presented in this paper. The 2D synthesis model is further evaluated in image denoising. Experimental results demonstrate our 2D synthesis sparse model outperforms the state-of-the-art 1D model in terms of both objective and subjective qualities. Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Jingdong Wang 0001 |
ICME | 3 |
| 2013 | Feature-based image set compressionabstractThe biggest challenge in image set compression is how to efficiently remove the set redundancy among images as well as the redundancy inside a single image. Different from all the previous schemes, in this paper we are the first to propose a generic image set compression scheme which removes the set redundancy based on local features in addition to luminance values. The SIFT (Scale Invariant Feature Transform) descriptor which characterizes an image region invariant to scale and rotation is utilized in our scheme to measure and further enhance the correlation among images. Given an image set, we build a minimal cost prediction structure according to the SIFT-based prediction measure between images. We also utilize a SIFT-based global transformation to enhance the correlation between two images by aligning them to each other in terms of both geometry and intensity. The set redundancy and image redundancy are both further reduced by block-based motion estimation and rate-distortion optimal mechanism proposed in HEVC. Experimental results show that our new feature based scheme always produces the best result regardless the image set's properties. Zhongbo Shi, Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 2 |
| 2013 | SIFT-based image super-resolutionabstractThis paper presents a new exemplar-based image super-resolution (SR) method in which we propose making use of scale invariant image features for high frequency (HF) approximation. We introduce the scale invariant feature transform (SIFT) descriptors in both building an exemplar dataset adaptively and producing the HF details with respect to the features of an input low resolution image. Given a large image database, we propose using the highly correlated images retrieved by SIFT descriptors for exemplar training rather than using a general set of images to increase the matching accuracy. Through building the training set of high resolution/low resolution exemplar pairs, the HF details for SR are retrieved from the training set by matching the SIFT features in a dense way. The flexibility as well as effectiveness of our SR approach is demonstrated at different magnification factors, e.g. 3 and 4. Experimental results show that our SIFT-based SR approach achieves enhanced high resolution images in terms of both objective and subjective qualities in comparison with the state-of-the-art exemplar-based methods. Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
ISCAS | 3 |
| 2013 | Multi-model prediction for image set compressionabstractThe key task in image set compression is how to efficiently remove set redundancy among images and within a single image. In this paper, we propose the first multi-model prediction (MoP) method for image set compression to significantly reduce inter image redundancy. Unlike the previous prediction methods, our MoP enhances the correlation between images using feature-based geometric multi-model fitting. Based on estimated geometric models, multiple deformed prediction images are generated to reduce geometric distortions in different image regions. The block-based adaptive motion compensation is then adopted to further eliminate local variances. Experimental results demonstrate the advantage of our approach, especially for images with complicated scenes and geometric relationships. Zhongbo Shi, Xiaoyan Sun 0001, Feng Wu 0001 |
VCIP | 2 |
| 2013 | Landmark Image Super-Resolution by Retrieving Web ImagesabstractThis paper proposes a new super-resolution (SR) scheme for landmark images by retrieving correlated web images. Using correlated web images significantly improves the exemplar-based SR. Given a low-resolution (LR) image, we extract local descriptors from its up-sampled version and bundle the descriptors according to their spatial relationship to retrieve correlated high-resolution (HR) images from the web. Though similar in content, the retrieved images are usually taken with different illumination, focal lengths, and shot perspectives, resulting in uncertainty for the HR detail approximation. To solve this problem, we first propose aligning these images to the up-sampled LR image through a global registration, which identifies the corresponding regions in these images and reduces the mismatching. Second, we propose a structure-aware matching criterion and adaptive block sizes to improve the mapping accuracy between LR and HR patches. Finally, these matched HR patches are blended together by solving an energy minimization problem to recover the desired HR image. Experimental results demonstrate that our SR scheme achieves significant improvement compared with four state-of-the-art schemes in terms of both subjective and objective qualities. Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Example-Based Super-Resolution With Soft Information and DecisionabstractThe one-to-one correspondence between co-occurrence image patches of two different resolutions is extensively used in example-based super-resolution (SR). Due to the dimensionality gap between low resolution (LR) and high resolution (HR) spaces, however, an LR patch may correspond to a number of HR patches in practice. This ambiguity is difficult to be overcome with examples representing a deterministic mapping. In this paper, we propose a statistical method for exploiting the one-to-many correspondence between LR and HR patches, which we call soft information and decision. Soft information means an LR patch is mapped to a pixel-wise distribution of all its possible HR counterparts, rather than a single or a limited set of HR candidates. Relying on the soft information, example-based SR is then regarded as an optimization problem to best preserve the local consistency in the recovered HR image. This problem is solved with an efficient message passing algorithm with a factor graph model. The final decision on the HR pixel value is made upon the maximum a posteriori estimation and is called a soft decision. Experimental results demonstrate the superiority of the proposed method compared with the state-of-the-art methods, in terms of both the subjective and objective quality of synthesized HR images. Zhiwei Xiong, Dong Xu 0001, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | Cloud-Based Image Coding for Mobile Devices - Toward Thousands to One CompressionabstractCurrent image coding schemes make it hard to utilize external images for compression even if highly correlated images can be found in the cloud. To solve this problem, we propose a method of cloud-based image coding that is different from current image coding even on the ground. It no longer compresses images pixel by pixel and instead tries to describe images and reconstruct them from a large-scale image database via the descriptions. First, we describe an input image based on its down-sampled version and local feature descriptors. The descriptors are used to retrieve highly correlated images in the cloud and identify corresponding patches. The down-sampled image serves as a target to stitch retrieved image patches together. Second, the down-sampled image is compressed using current image coding. The feature vectors of local descriptors are predicted by the corresponding vectors extracted in the decoded down-sampled image. The predicted residual vectors are compressed by transform, quantization, and entropy coding. The experimental results show that the visual quality of reconstructed images is significantly better than that of intra-frame coding in HEVC and JPEG at thousands to one compression . Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Spatially Scalable Video Coding for HEVCabstractSpatially scalable video coding (SSVC in short) provides an efficient way to deliver one video at different resolutions. Based on the development of emerging High Efficiency Video Coding (HEVC), we propose a SSVC scheme to provide both single-loop and multi-loop solutions by enabling different inter-layer prediction mechanisms. Specifically, there are three basic inter-layer prediction modes, inter-layer intra, motion and residual prediction, which are first investigated for enhancement layer coding. These three modes provide a basic single-loop SSVC solution for low complexity applications. Beside the correlations explored in single-loop solution, we employ an extra patch learning-based prediction mode named P-mode to further improve coding performance. In P-mode, the temporal and spatial correlations of the reference frames are explored simultaneously by the visual patch-based learning and mapping at pixel level. These correlations help enhance the accuracy of the inter-layer frame prediction from the base-layer reconstruction within a multi-loop structure. The experimental results show the effectiveness of the proposed SSVC scheme compared with simulcast case. Zhongbo Shi, Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 2 |
| 2012 | SIFT-Based Image CompressionabstractThis paper proposes a novel image compression scheme based on the local feature descriptor - Scale Invariant Feature Transform (SIFT). The SIFT descriptor characterizes an image region invariantly to scale and rotation. It is used widely in image retrieval. By using SIFT descriptors, our compression scheme is able to make use of external image contents to reduce visual redundancy among images. The proposed encoder compresses an input image by SIFT descriptors rather than pixel values. It separates the SIFT descriptors of the image into two groups, a visual description which is a significantly sub sampled image with key SIFT descriptors embedded and a set of differential SIFT descriptors, to reduce the coding bits. The corresponding decoder generates the SIFT descriptors from the visual description and the differential set. The SIFT descriptors are used in our SIFT-based matching to retrieve the candidate predictive patches from a large image dataset. These candidate patches are then integrated into the visual description, presenting the final reconstructed images. Our preliminary but promising results demonstrate the effectiveness of our proposed image coding scheme towards perceptual quality. Our proposed image compression scheme provides a feasible approach to make use of the visual correlation among images. Huanjing Yue, Xiaoyan Sun 0001, Feng Wu 0001, Jing-Yu Yang 0002 |
ICME | 2 |
| 2012 | IMShare: instantly sharing your mobile landmark images by search-based reconstructionabstractInstantly sharing captured landmark images is becoming fashionable, much like when you write a blog or chat with friends by mobile phone. However, real-time transmission of high-resolution images poses a significant challenge to contemporary mobile networks. Either long delays in transmission or largely reduced image resolution can lead to bad user experience. In this paper, we propose a novel mobile-cloud scheme IMShare to enable instant sharing of high-resolution images. On the mobile side, high-resolution images are described by their thumbnails and SIFT (Scale-Invariant Feature Transform) descriptors. After compression, data sent by mobile phones can be reduced to an average of 2.6 kilobytes (KB) per mega pixel. On the cloud side, high-resolution images are reproduced from a large-scale image database by retrieving partial duplicate images by SIFT descriptors and stitching corresponding image patches together under the guidance of the thumbnails. IMShare is the first scheme to demonstrate that not only visually pleasant images can be reconstructed using this mobile-cloud method but also the reconstruction can be done in seconds using parallel computing. Our user study of a half million images in a database shows that the proposed IMShare significantly outperforms the current method on subjective quality. Lican Dai, Huanjing Yue, Xiaoyan Sun 0001, Feng Wu 0001 |
ACM Multimedia | 3 |
| 2012 | Inpainting with image patches for compression
Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | Content-adaptive deblocking for high efficiency video coding
Zhiwei Xiong, Xiaoyan Sun 0001, Jizheng Xu, Feng Wu 0001 |
Signal Process. Image Commun. | 2 |
| 2012 | Spatially Scalable Video Coding For HEVCabstractSpatially scalable video coding (SSVC) provides an efficient way to transmit one video at different resolutions. Based on the emerging High Efficiency Video Coding (HEVC), we propose an SSVC scheme to support both single-loop (SL) and multiloop (ML) solutions by enabling different interlayer prediction mechanisms. Specifically, we employ two interlayer prediction modes: quadtree-based prediction mode (Q-mode) and learning-based prediction mode (L-mode). The Q-mode is investigated to exploit the interlayer redundancy based on the quadtree coding structure of HEVC. Due to the high correlation between layers, Q-mode utilizes the coded information from the base layer quadtree, including coding unit split, prediction unit partition, motion information, and partial texture information of transform unit, to predict the enhancement layer quadtree. By enabling Q-mode, we provide a basic SL solution for low complexity applications. Besides the correlation explored in Q-mode, we employ an extra L-mode to further improve the coding performance. In L-mode, the temporal-spatial correlation is exploited simultaneously by visual patch-based learning and mapping at pixel level. This helps us achieve more accurate prediction signals based on the coarse base layer reconstruction within an ML structure. Experimental results show the effectiveness of our SSVC scheme compared with the simulcast case and other HEVC-based SSVC schemes. Zhongbo Shi, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Adaptive patch matching for motion compensated predictionabstractMotion compensated prediction (MCP) plays an important role in video coding due to its great capability of reducing temporal redundancy. In this paper, we propose a new MCP scheme by adaptive patch matching with the full use of the reconstructed pixels surrounding the current block (referred to as the template inside the patch) aiming at achieving a more accurate prediction than conventional MCP. The proposed scheme not only takes advantage of the temporal correlation but also efficiently exploits the spatial correlation between the current block and its template inside the patch. An adaptive linear combination of the current block and its template in motion estimation is designed to generate an optimal prediction while maintaining the local variation of the current block. Accordingly, a modification of the rate-distortion criterion is introduced to select the combined prediction. Experimental results show that our proposed APM achieves improved coding performance compared with H.264/AVC. Tianmi Chen, Xiaoyan Sun 0001, Feng Wu 0001, Guangming Shi |
ISCAS | 2 |
| 2011 | CGS quality scalability for HEVCabstractScalable video coding provides an efficient way to serve video contents at different quality levels. Based on the development of emerging High Efficiency Video Coding (HEVC), we propose two coarse granular scalable (CGS) video coding schemes here. In scheme A, we present a multi-loop solution in which the fully reconstructed base pictures are utilized in the enhancement layer prediction. By inserting the reconstructed base picture (BP) into the list of reference pictures of the collocated enhancement layer frame, we enable the coarse granular quality scalability of HEVC with very limited changes. On the other hand, scheme B supports single loop decoding. It contains three inter-layer predictions similar to the scalable extension of H.264/AVC. Compared to scheme A, it decreases the decoding complexity by avoiding the motion compensation, deblocking filtering (DF) and adaptive loop filtering (ALF) in the base layer. The effectiveness of our proposed two coding schemes is evaluated by comparing with single-layer coding and simulcast. Zhongbo Shi, Xiaoyan Sun 0001, Jizheng Xu |
MMSP | 2 |
| 2010 | Predictive patch matching for inter-frame codingabstractInter prediction is an important component for video coding which exploits the temporal correlation between frames and significantly reduces the redundancy in video sequences. In this paper, we propose a predictive patch matching for inter prediction based on template matching prediction. Besides the surrounding reconstructed pixels which form the template in template matching prediction, our proposed patch matching is able to utilize the predicted pixels generated by the traditional motion prediction. A linear combination of the reconstructed template and the predicted pixels permits to synthesize a prediction while maintaining the local variations of the target block. Furthermore, a mode selection mechanism is introduced to adaptively select the predictive patch matching at sub-block level. Experimental results demonstrate the effectiveness of our proposed predictive patch matching. Constant coding gain can be achieved by our scheme at both low and high bit rates Tianmi Chen, Xiaoyan Sun 0001, Feng Wu 0001 |
VCIP | 2 |
| 2010 | Block-Based Image Compression With Parameter-Assistant InpaintingabstractThis correspondence presents an image compression approach that integrates our proposed parameter-assistant inpainting (PAI) to exploit visual redundancy in color images. In this scheme, we study different distributions of image regions and represent them with a model class. Based on that, an input image at the encoder side is divided into featured and non-featured regions at block level. The featured blocks fitting the predefined model class are coded by a few parameters, whereas the non-featured blocks are coded traditionally. At the decoder side, the featured regions are restored through PAI relying on both delivered parameters and surrounding information. Experimental results show that our method outperforms JPEG in featured regions by an average bit-rate saving of 76% at similar perceptual quality levels. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2010 | Robust Web Image/Video Super-ResolutionabstractThis paper proposes a robust single-image super-resolution method for enlarging low quality web image/video degraded by downsampling and compression. To simultaneously improve the resolution and perceptual quality of such web image/video, we bring forward a practical solution which combines adaptive regularization and learning-based super-resolution. The contribution of this work is twofold. First, we propose to analyze the image energy change characteristics during the iterative regularization process, i.e., the energy change ratio between primitive (e.g., edges, ridges and corners) and nonprimitive fields. Based on the revealed convergence property of the energy change ratio, appropriate regularization strength can then be determined to well balance compression artifacts removal and primitive components preservation. Second, we verify that this adaptive regularization can steadily and greatly improve the pair matching accuracy in learning-based super-resolution. Consequently, their combination effectively eliminates the quantization noise and meanwhile faithfully compensates the missing high-frequency details, yielding robust super-resolution performance in the compression scenario. Experimental results demonstrate that our solution produces visually pleasing enlargements for various web images/videos. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | Image hallucination with feature enhancementabstractExample-based super-resolution recovers missing high frequencies in a magnified image by learning the correspondence between co-occurrence examples at two different resolution levels. As high-resolution examples usually contain more details and are of higher dimensionality in comparison with low-resolution ones, the mapping from low-resolution to high-resolution is an ill-posed problem. Rather than imposing more complicated mapping constraints, we propose to improve the mapping accuracy by enhancing low-resolution examples in terms of mapped features, e.g., derivatives and primitives. A feature enhancement method is presented through a combination of interpolation with prefiltering and non-blind sparse prior deblurring. By enhancing low-resolution examples, unique feature information carried by high-resolution examples is decreased. This regularization reduces the intrinsic dimensionality disparity between two different resolution examples and thus improves the feature mapping accuracy. Experiments demonstrate our super-resolution scheme with feature enhancement produces high quality results both perceptually and quantitatively. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
CVPR | 2 |
| 2009 | Improving Inverse Wavelet Transform by Compressive Sensing Decoding with DeconvolutionabstractIn this paper we propose an alternative decoding method for inverse wavelet transform when only partial coefficients are available. We have been inspired by the recently developed compressive sensing (CS) decoding, which is capable in recovering sparse signals from a few linear and non-adaptive measurements. Let x be a sparse signal with N entries and only K out of them are non-zero, and y be its approximation coefficients. Classic CS decoding such as l1-minimization can be applied to decode x from y, and it indeed provides better reconstruction of sparse signals than direct inverse transform, as demonstrated by our simulation results. When coefficients have been quantized, the performance of CS decoding decreases more severely compared with direct inverse transform, but still better than the latter once the signal is sparse enough. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
DCC | 2 |
| 2009 | Classified patch learning for spatially scalable video codingabstractThis paper proposes an advanced spatially scalable video coding approach that exploits the inter layer correlation between different resolution layers by classified patch learning. The novelty of our proposed scheme is twofold. First, the correlation between low and high resolution frames is explored at patch level with regard to image features. Patches extracted from the previous coded frame are classified into structural and textural sets according to the gradient information. Then the inter layer correlation is separately studied for the two sets, resulting in two databases containing pairs of patches at different resolutions. Second, our proposed patch-based compensation manages to simultaneously exploit the spatial and temporal redundancies without overhead bit for motion. Based on the two databases, a high resolution prediction is derived from the current low resolution reconstruction at structural and textural regions, respectively. Experimental results show that our proposed approach improves the performance of H.264/MPEG spatially scalable coding up to 1.9 dB and significantly enhances the subjective quality, especially at low bit rates. Xiaoyan Sun 0001, Feng Wu 0001 |
ICIP | 1 |
| 2009 | Web cartoon video hallucinationabstractThis paper addresses the super-resolution problem for low quality cartoon videos widely distributed on the web, which are generated by downsampling and compression from the sources. To effectively eliminate the compression artifacts and meanwhile preserve the visually salient primitive components (e.g., edges, ridges and corners), we propose an adaptive regularization method depending on the degradation grade of each frame, followed by learning-based pair matching to further enhance the primitives in the upsampled frames. In addition, temporal consistency is considered a directive constraint in both the regularization and enhancement processes. Experimental results demonstrate our solution achieves a good balance between artifacts removal and primitive enhancement, providing perceptually high quality super-resolution results for various web cartoon videos. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
ICIP | 2 |
| 2009 | Fractional compensation for spatial scalable video codingabstractThis paper proposes a novel fractional compensation approach for spatial scalable video coding. It simultaneously exploits inter layer correlation and intra layer correlation by learning-based mapping. Instead of using an enhancement layer reconstruction as an entire reference, a set of reference pairs are generated from high-frequency components of both base layer and enhancement layer reconstructions at previous frame. The reference set, which consists of low-resolution and high-resolution patches, can be generated in both encoder and decoder by on-line learning. During the encoding of enhancement layer, a prediction is first gotten from base layer, from which low-resolution patches are extracted. These patches are then used as indices to find the matched high-resolution patches from the reference set. Finally, the prediction enhanced by the high-resolution patches is used for coding. The proposed approach does not need any motion bits. With our proposed FC approach, the performance of H.264 SVC can be improved up to 2.4 dB in spatial scalable coding. Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 1 |
| 2009 | Fast directional image interpolation with difference projectionabstractThis paper presents a new directional image interpolator, aiming to increase image resolution with high perceptual quality and low computational complexity. In our method, missing pixels in a magnified image are generated through linear interpolation on certain fixed supports to facilitate fast implementation, while local directional features are imposed on the adaptive interpolation weights which are determined by the gradients diffused from the low resolution image. Afterwards, a novel difference projection strategy is proposed to enforce the continuity of the magnified image by reusing the directional interpolator. Experimental results show that our method outperforms conventional bicubic and some existing adaptive interpolators, in terms of both the perceptual and quantitative quality. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 3 |
| 2008 | Intra Prediction via Edge-Based InpaintingabstractWe investigate the usage of edge-based inpainting as an intra prediction method in block-based image compression. The joint utilization of edge information and the well-known Laplace equation yields a simple and effective inpainting algorithm. As for intra prediction, the edge-based inpainting is a uniform solution, yet adaptive to local image features. During the integration of edge-based inpainting into a block-based coding scheme, edge extraction and coding are jointly considered to achieve the rate-distortion optimization. Our proposed schemes are compared with JPEG2000, and experimental results demonstrate that both PSNR gain and visible quality improvement are achieved. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
DCC | 2 |
| 2008 | Image Compression by Visual Pattern Vector Quantization (VPVQ)abstractThis paper proposes a new image compression scheme by introducing visual patterns to nonlinear interpolative vector quantization (IVQ). Input images are first distorted by a generic down-sampling so that some details are removed before compression. Then, the distorted images are compressed lossly by traditional image coding scheme and transmitted to the decoder. In the decoder side, VQ indices are extracted from the decoded images to reproduce the removed details from a pre-trained codebook. One of main contributions in this paper is, we introduce visual patterns on designing the codebook, where only removed details that contain visual patterns and their original counterparts as pairs are trained. Experimental results show: (1) visual pattern blocks are easy to form clusters than original blocks; (2) the proposed scheme achieves much better performance over JPEG in terms of visual quality and PSNR. Feng Wu 0001, Xiaoyan Sun 0001 |
DCC | 2 |
| 2008 | Manipulating image patches for compressionabstractWe consider how to exploit the correlation in image for compression by virtue of studying image patches in a non-parametric manner. Instead of extracting and recording parameters, our approach directly operates on image patches. The basic assumption is that a subset of image patches can be well inferred from the others; therefore, they can be removed at encoder only to be restored at decoder. Meanwhile, assistant information is transmitted for the restoration, which actually encodes the similarity between removed and preserved patches. The entire scheme is built upon an optimization framework, which is decoupled and solved accordingly. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 2 |
| 2008 | Super-resolution for low quality thumbnail imagesabstractThis paper proposes a single-image super-resolution scheme for enlarging low quality thumbnail images widely distributed on the web, which are often generated by downsampling plus compression. To obtain visually pleasurable high-resolution versions for this kind of low-resolution images, we first adopt a PDE-based image regularization technique to alleviate the compression noise in the distorted thumbnails, and then use learning-based pair matching to further enhance the high-frequency details in the upsampled images. Experimental results show that our solution achieves better visual quality for both offline and online test images, compared with traditional methods. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 2 |
| 2008 | Video coding with spatio-temporal texture synthesis and edge-based inpaintingabstractThis paper proposes a video coding scheme, in which textural and structural regions are selectively removed in the encoder, and restored in the decoder by spatio-temporal texture synthesis and edge-based inpainting. In the proposed scheme, two types of regions are classified based on two motion models: local motion and global motion. In local motion regions, conventional block-based motion estimation is employed for region removal and spatio-temporal texture synthesis is applied for recovery of the removed regions. In global motion regions, edge-based image inpainting is utilized to recover removed regions, and sprite generation is used as an auxiliary tool to keep temporal consistency. In the proposed scheme, both structures and textures are handled and some kinds of assistant information which can guide restoration are extracted and coded. This approach is blockbased and thus is flexible and generic to be implemented into standard-compliant video coding schemes. It has been implemented into H.264/AVC and achieves up to 35% bitrate saving at similar visual quality levels compared with H.264/AVC without our approach. Chunbo Zhu, Xiaoyan Sun 0001, Feng Wu 0001, Houqiang Li |
ICME | 2 |
| 2008 | Fast H.264/MPEG-4 AVC Transcoding Using Power-Spectrum Based Rate-Distortion OptimizationabstractSince variable block-size motion compensation (MC) and rate-distortion optimization (RDO) techniques are adopted in H.264/MPEG-4 AVC, modes and motion vectors (MVs) in input stream can no longer be reused equivalently efficient over a wide range of bit rate in transcoded streams. This paper proposes a new RDO model to maintain good coding efficiency and greatly reduce computation of the H.264/MPEG-4 AVC transcoding, in which the distortion caused by motion and mode changes is not calculated directly from the sum of absolute difference (SAD) or the sum of square difference (SSD) between source signals and interpolated prediction signals. Instead, distortion is directly estimated from MV variation and the power spectrum (PS) of the prediction signal generated from input stream. The proposed RDO model can be applied to both the pixel-domain transcoding and the transform-domain transcoding even when coded signals are not reconstructed at all. Furthermore, the techniques as to derive the Lagrangian multiplier in the proposed model are developed in respective pixel- and transform-domains. Additionally, we propose an H.264/MPEG-4 transcoding scheme that demonstrates the advantage of the proposed RDO model in terms of peak signal-to-noise ratio and transcoding speed, in which P-pictures are transcoded in the pixel domain for achieving reconstructed high quality and B-pictures are transcoded in the transform domain for high-transcoding speed. Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Edge-Oriented Uniform Intra PredictionabstractWe propose an intra prediction solution to block-based image compression. In order to adapt to local image features during intra prediction, we consider the distinct image singularities within the model of piece-wise smooth functions. With such singularities, i.e., edges in this paper, intra prediction can be performed by solving Laplace equations. Moreover, since edges exhibit spatial correlations, we design a rate-distortion optimized method for edge extraction and edge coding. Our edge-oriented intra prediction thus consists of the prediction of smooth regions as well as the prediction of edges. We compare our intra prediction with that in H.264 and achieve superior performance. Our intra prediction can also be integrated into a block-based image coding scheme, which is comparable to JPEG2000 in terms of objective quality. An important advantage of our intra prediction is the improvement in visual quality at low bit-rate due to the preservation of edges. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001, Ya-Qin Zhang |
IEEE Trans. Image Process. | 2 |
| 2007 | Incorporating Primal Sketch Based Learning Into Low Bit-Rate Image CompressionabstractThis paper proposes an image compression approach, in which we incorporate primal sketch based learning into the mainstream image compression framework. The key idea of our approach is to use primal sketch information to enhance the quality of distorted images. With this method, we only encode the down-sampled image and use the primal sketch based learning to recover the high frequency information which has been removed by down-sampling. Experimental results demonstrate that our scheme achieves better objective visual quality as well as subjective quality compared with JPEG2000 at the same bit-rates. Xiaoyan Sun 0001, Hongkai Xiong, Feng Wu 0001 |
ICIP (3) | 2 |
| 2007 | Image Coding with Parameter-Assistant InpaintingabstractThis paper carves out an image compression approach that integrates our parameter-assistant inpainting (PAI) technique to exploit the visual redundancy inherent in color-gradation image regions. In our scheme, an input image is first classified at block level according to the degree of edge content as well as chromatic variation in each block. An exemplar selection approach is then adopted to skip a majority of the gradation blocks during encoding. Only their positions and certain parameters extracted for condensed description are encoded along with the reserved blocks. At the decoder side, the skipped regions are recovered through image inpainting, relying on both the delivered parameters and reserved regions. Experimental results show that our proposed method outperforms baseline JPEG at color-gradation regions by nearly 80% bits-saving, at similar visual quality levels. Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001 |
ICIP (2) | 2 |
| 2007 | Edge-Based Inpainting and Texture Synthesis for Image CompressionabstractTowards visual quality rather than pixel-wise fidelity, we propose an image coding scheme integrated with edge-based in-painting and texture synthesis. In this scheme, an original image is analyzed at encoder side so that some blocks are removed during encoding. The edges related to these removed blocks will be compressed and transmitted. At decoder side, we propose an image restoration method, which consists of edge-based in-painting and texture synthesis, in order to fully utilize the transmitted edges and naturally restore the removed blocks. Experimental results show that our scheme can achieve up to 32% bit-rate saving at similar visual quality levels, compared with H.264/AVC intra coding. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001 |
ICME | 2 |
| 2007 | Video Coding with Spatio-Temporal Texture SynthesisabstractThis paper presents a video coding scheme in which some texture regions are selectively removed at the encoder and recovered by synthesis at the decoder. We present region removal utilizing conventional block-based motion information rather than global motion field. Removed regions including their motion information are not coded at the encoder. We propose spatio-temporal patch-searching in texture synthesis at the decoder to recover the removed regions. Our approach is not content based and is flexible and generic to be implemented. The scheme has been integrated into H.264/AVC and achieves up to 38.8% bitrate saving at similar visual quality levels compared with H.264/AVC. Chunbo Zhu, Xiaoyan Sun 0001, Feng Wu 0001, Houqiang Li |
ICME | 2 |
| 2007 | Image Compression With Edge-Based InpaintingabstractIn this paper, image compression utilizing visual redundancy is investigated. Inspired by recent advancements in image inpainting techniques, we propose an image compression framework towards visual quality rather than pixel-wise fidelity. In this framework, an original image is analyzed at the encoder side so that portions of the image are intentionally and automatically skipped. Instead, some information is extracted from these skipped regions and delivered to the decoder as assistant information in the compressed fashion. The delivered assistant information plays a key role in the proposed framework because it guides image inpainting to accurately restore these regions at the decoder side. Moreover, to fully take advantage of the assistant information, a compression-oriented edge-based inpainting algorithm is proposed for image restoration, integrating pixel-wise structure propagation and patch-wise texture synthesis. We also construct a practical system to verify the effectiveness of the compression approach in which edge map serves as assistant information and the edge extraction and region removal approaches are developed accordingly. Evaluations have been made in comparison with baseline JPEG and standard MPEG-4 AVC/H.264 intra-picture coding. Experimental results show that our system achieves up to 44% and 33% bits-savings, respectively, at similar visual quality levels. Our proposed framework is a promising exploration towards future image and video compression. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Transcoding to FGS Streams from H.264/AVC Hierarchical B-PicturesabstractThis paper presents a transcoder which transcodes to FGS streams from H.264/AVC hierarchical B-pictures. First, the DCT-domain architecture is designed for fast FGS transcoding. Then, we propose a mode decision method in DCT domain to achieve a trade-off between the performances at low bit-rate and high bit-rate. Experimental results demonstrated that our method can improve the coding performance up to 1 dB at low rate and only lose at worst 0.5 dB at high rate. Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001, Houqiang Li, Shipeng Li 0001 |
ICIP | 2 |
| 2006 | A Fast Downsizing Video Transcoder for H.264/AVC with Rate-Distortion Optimal Mode DecisionabstractThis paper focuses on the mode decision and motion selection problem when H.264/AVC video streams are transcoded in spatial resolution. A fast downsizing transcoding scheme is developed in which a new rate-distortion (R-D) optimal mode decision mechanism is presented for high speed transcoding as well as high coding efficiency. A model for estimating relative prediction errors is applied in this paper, which is free from computation of interpolation and SAD/SSD computation. Based on the selected model, a motion refinement within a distance of 1 pixel is performed after mode decision. Experimental results demonstrate that our method can significantly speed up the spatial resolution reduction process, while achieving high coding efficiency Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001, Houqiang Li, Shipeng Li 0001 |
ICME | 2 |
| 2006 | Off-Line Motion Description for Fast Video Stream Generation in MPEG-4 AVC/H.264abstractThe rate-distortion optimal mode decision as well as motion estimation adopted in H.264 brings a big challenge to real-time encoding and transcoding due to the high computation complexity. In this paper, we propose a hierarchical motion description model to present the motion data of each macroblock (MB) from coarsely to finely. A preprocessing approach is developed to estimate the motion data for each MB at each quality level with regard to its reference quality, its adjacent MBs and the target bit-rate. The resulting motion data can be coded and stored as metadata in a media file or a stream. Moreover, we propose a method to readily extract the specific motion data from the model for each MB at given bit-rates. Experimental results have shown the effectiveness of our proposed motion description model in terms of coding efficiency as well as fast bit-rate adaptation in comparison with that of H.264 Yi Wang 0037, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Houqiang Li, Zhengkai Liu |
ICME | 2 |
| 2006 | Rate-distortion optimization for fast hierarchical B-picture transcodingabstractAn efficient rate-distortion (R-D) optimal method for transcoding hierarchical B-pictures is proposed in this paper. A new R-D model is presented for fast transcoding hierarchical B-pictures in DCT domain, in which the total R-D optimization problem is adjusted to motion R-D optimization and texture R-D optimization separately. Accordingly, a mechanism for fast mode decision is proposed to enable the mode and motion adjustment in hierarchical B-picture DCT-domain transcoding. Experimental results show that the proposed transcoding scheme with the new R-D optimization can achieve 4-dB average PNSR improvement at low bit-rate, and similar performance at high bit-rate, compared to the transcoding by reusing the input motion information. Moreover, as the transcoding is performed in DCT domain, it is fast and simple enough for many real applications. Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001 |
ISCAS | 2 |
| 2006 | Image compression with structure-aware inpaintingabstractThis paper carves out a way to image compression that is motivated by the recent advancement in image inpainting. An image coding approach is proposed in which a number of regions of the input image are skipped at the encoder and are recovered through the inpainting process at the decoder. Furthermore, a structure-aware inpainting (SAI) method is developed to restore the skipped structural regions by taking advantage of the available portion of the decoded image. A binary structure map is extracted and compressed into the generated bit-stream to indicate the skipped regions with salient structures. By making use of the decoded texture information together with the structure map, the SAI method can recover the skipped structural regions as well as the non-structural ones effectively at the decoder. Compared with JPEG, our proposed image compression scheme allows smaller file, with the potential of up to 50% bit-saving capability, at similar visual quality levels Chen Wang 0053, Xiaoyan Sun 0001, Feng Wu 0001, Hongkai Xiong |
ISCAS | 2 |
| 2006 | Drift-free switching of compressed video bitstreams at predictive framesabstractTwo schemes are proposed to efficiently compress video contents into bitstreams that support drift-free switching at predictive frames. They are inspired by the original SP coding scheme presented in the early H.26L. First, we propose a Flex SP coding scheme in which the prediction signal of the SP frame is directly subtracted from the input without quantization and de-quantization. The decoded video quality of the Flex SP scheme is significantly improved when additional inverse discrete cosine transform (DCT) and post-filter are provided. Then, the Hybrid SP scheme is presented to further improve the quality of the display image, as well as the reconstructed reference, by defining two coding modes for each DCT coefficient. Moreover, a rate-distortion algorithm is proposed to determine the coding mode for each coefficient. The bitstreams generated by the two proposed schemes can be decoded successfully by a decoder that complies with MPEG-4 AVC/H.264. In addition, we also investigate how to choose the quantization parameters for switching. An empirical method is proposed to achieve a good tradeoff between high coding efficiency of SP frames and small size of switching bits. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Guobin Shen, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Spatio-temporal video error concealment using priority-ranked region-matchingabstractWhen transmitted over error-prone networks, compressed video sequences may be received with errors. In this paper, we propose a priority-ranked region-matching algorithm to recover the "lost" area of the decoded frames, in which both temporal and spatial correlations of the video sequence are exploited. In the proposed scheme, we first calculate the priorities of all edge pixels of the "lost" area and generate a priority-ranked region group. Then according to their priorities, the regions in the group will search their best matching regions temporally and spatially. Finally, the "lost" area is recovered progressively by the corresponding pixels in the matching regions. Experimental results show that the proposed scheme achieves higher PSNR as well as better video quality in comparison with the method adopted in H.264. Yan Chen 0007, Xiaoyan Sun 0001, Feng Wu 0001, Zhengkai Liu, Shipeng Li 0001 |
ICIP (2) | 2 |
| 2005 | H.264-compatible spatially scalable video coding with in-band predictionabstractIn this paper, a H.264 compatible spatially scalable video coding method with in-band prediction is proposed which taking advantages from both the high coding efficiency of H.264 coding scheme and the attractive performance of in-band overcomplete discrete wavelet transform (ODWT) in wavelet-domain motion estimation and motion compensation. Four MV prediction modes are proposed for INTER prediction of high frequency subbands. The intra prediction modes of H.264 are also simplified for each high band according to the directional features inherited inside. Finally, a H.264 compatible scheme based on one of the MV prediction modes is presented to provide better tradeoff among standard compatibility, low complexity and high performance. Xin Jin 0014, Xiaoyan Sun 0001, Feng Wu 0001, Guangxi Zhu, Shipeng Li 0001 |
ICIP (1) | 2 |
| 2004 | Weighted motion estimation for efficiently coding scene transition videoabstractScene transition video has brought great challenges to current video coding methods because the traditional motion model of block displacement cannot efficiently represent transition motion. The paper first analyzes the features of static transitions. By incorporating the transition filter information into the coding scheme, the weighted motion estimation (WTME) technique is proposed to get accurate motion parameters for the transition video, thereby efficiently compensating both normal and transition motions among frames. Experimental results show that the proposed technique can significantly improve the coding performance of H.264 by up to 2.0 dB while coding scene transition video. Xiaoyan Sun 0001, Hong Bao, Shipeng Li 0001 |
ICASSP (3) | 2 |
| 2004 | Variable block-size transform and entropy coding at the enhancement layer of FGSabstractThis paper proposes a variable block-size transform and context-based entropy coding techniques for the enhancement layer of FGS (fine granularity scalable) video coding. First, the variable block-size transform is introduced into the enhancement layer to improve the performance of FGS in terms of both visual quality and PSNR. Different from that used in the traditional single layer coding, an R-D selection algorithm is proposed to optimally decide the transform size of each block, under consideration of consistent performance at a range of bit rates. Furthermore, to fully take advantage of the characteristics and correlations of symbols coded in the FGS enhancement layer, different context models are designed for the arithmetic coding according to symbol type and transform size. Experimental results show that the coding efficiency of FGS can be increased by 0.2-0.90 dB with the proposed techniques. Jungong Han, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Zhaoyang Lu |
ICIP | 2 |
| 2004 | Flexible p-picture (FLEXP) coding for tue efficient fine-granular scalabilitv (FGS)
Xiaoyan Sun 0001, Feng Wu 0001, Hong Bao, Shipeng Li 0001 |
ICIP | 2 |
| 2004 | Seamless switching of scalable video bitstreams for efficient streamingabstractEfficient adaptation to channel bandwidth is broadly required for effective streaming video over the Internet. To address this requirement, a novel seamless switching scheme among scalable video bitstreams is proposed in this paper. It can significantly improve the performance of video streaming over a broad range of bit rates by fully taking advantage of both the high coding efficiency of nonscalable bitstreams and the flexibility of scalable bitstreams, where small channel bandwidth fluctuations are accommodated by the scalability of a single scalable bitstream, whereas large channel bandwidth fluctuations are tolerated by flexible switching between different scalable bitstreams. Two main techniques for switching between video bitstreams are proposed. Firstly, a novel coding scheme is proposed to enable drift-free switching at any frame from the current scalable bitstream to one operated at lower rates without sending any overhead bits. Secondly, a switching-frame coding scheme is proposed to greatly reduce the number of extra bits needed for switching from the current scalable bitstream to one operated at higher rates. Compared with existing approaches, such as switching between nonscalable bitstreams and streaming with a single scalable bitstream, our experimental results clearly show that the proposed scheme brings higher efficiency and more flexibility in video streaming. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 1 |
| 2003 | The improved SP frame coding technique for the JVT standardabstractAn efficient and flexible coding technique is proposed in this paper inspired by the SP frame in the H.26L standard, which can achieve a drift-free bitstream switching at the predicted frame. The proposed scheme improves the coding efficiency of the SP frames in the H.26L standard by limiting the mismatch between the references for the prediction and reconstruction with two DCT coefficient coding modes and the rate-distortion optimization. Furthermore, the proposed scheme allows independent quantization parameters for up-switching and down-switching bitstreams. It further reduces the switching bitstream size while keeping the coding efficiency of the normal bitstreams. More rapid and frequent down-switching than up-switching and much smaller size of down-switching bitstream can be achieved with the proposed SP technique. These are very desirable features for any TCP-friendly protocols. Compared with the original SP method for H.26L, the proposed SP method improves the coding efficiency up to 1.0 dB. This SP technique has been officially accepted by the JVT standard. Xiaoyan Sun 0001, Shipeng Li 0001, Feng Wu 0001, Guobin Shen, Wen Gao 0001 |
ICIP (3) | 1 |
| 2002 | Efficient and universal scalable video codingabstractThis paper proposes a unified efficient and universal scalable video coding framework that supports different scalabilities, such as fine granularity quality, temporal, spatial and complexity scalabilities. The proposed framework is established upon the recent studies in fine granularity scalable (FGS) video coding. It contains two key points. Firstly, in order to improve the coding efficiency of the proposed framework, more than one motion compensation loop is used. Since high quality references are introduced into the enhancement layer coding, the proposed framework can efficiently compress different-resolution video at different layers for the purpose of the complexity and spatial scalability. Secondly, the drifting reduction techniques are studied in this paper. This helps the proposed framework to maintain good performance at lower enhancement bit rates. By defining coding modes, a macroblock level control mechanism is developed to achieve a better trade-off between low drifting errors and high coding efficiency. Feng Wu 0001, Shipeng Li 0001, Xiaoyan Sun 0001, Ya-Qin Zhang |
ICIP (2) | 4 |
| 2002 | Efficient and flexible drift-free video bitstream switching at predictive framesabstractWe propose an efficient and flexible coding scheme inspired by the SP picture technique in H.26L TML; it can achieve drift-free bitstream switching at predictive frames. Firstly, the proposed scheme improves the coding efficiency of the SP frames in H.26L TML by (1) reducing the number of quantization modules in the encoding path; (2) eliminating the mismatch between references for the prediction and the reconstruction; (3) outputting a high quality image for display purpose before the quantization step in the reconstruction loop. Secondly, the proposed scheme allows independent quantization parameters for up-switching and down-switching bitstreams. It can further reduce the switching bitstream size while keeping the coding efficiency of the normal bitstreams. It allows more rapid and frequent down-switching than up-switching. Furthermore, the size of the down-switching bitstream can be much smaller than that of the up-switching one. This is a very desirable feature for any TCP-friendly protocols currently used in most existing streaming systems. Xiaoyan Sun 0001, Shipeng Li 0001, Feng Wu 0001, Guobin Shen, Wen Gao 0001 |
ICME (1) | 1 |
| 2001 | Macroblock-based progressive fine granularity scalable (PFGS) video coding with flexible temporal-SNR scalablilitiesabstractWe proposed a flexible and efficient architecture for scalable video coding, namely, the macroblock (MB)-based progressive fine granularity scalable video coding with temporal-SNR scalabilities (PFGST). The proposed architecture can provide not only much improved coding efficiency but also simultaneous SNR scalability and temporal scalability. Building upon the original frame-based progressive fine granularity scalable (PFGS) coding approach, the MB-based PFGS scheme is first proposed. Three INTER modes and the corresponding mode selection mechanism are presented for coding the SNR enhancement MBs in order to make a good trade-off between low drifting errors and high compression efficiency. Furthermore, temporal scalability is introduced into the MB-based PFGS, which forms the MB-based PFGST scheme. Two coding modes are proposed for coding the temporal enhancement MBs. Since it would not cause any error propagation if using the high quality reference in the temporal enhancement MB coding, the coding efficiency of the PFGST is highly improved by always choosing the most suitable reference for the temporal scalable coding. Experimental results show that the MB-based PFGST video coding scheme can significantly improve the coding efficiency up to 2.8 dB compared with the FGST scheme adopted in MPEG-4, while supporting full SNR, full temporal, and hybrid SNR-temporal scalabilities according to the different requirements from the channels, the clients or the servers. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang |
ICIP (2) | 1 |
| 2001 | Macroblock-Based Progressive Fine Granularity Scalable Video CodingabstractIn this paper, we proposed a flexible and efficient architecture for scalable video coding, namely, the macroblock (MB)-based progressive fine granularity scalable video coding with temporal-SNR scalabilities (PFGST in short). The proposed architecture can provide not only much improved coding efficiency but also simultaneous SNR scalability and temporal scalability. Building upon the original frame-based progressive fine granularity scalable (PFGS) coding approach, the MB-based PFGS scheme is first proposed. Three INTER modes and the corresponding mode selection mechanism are presented for coding the SNR enhancement MBs in order to make a good trade-off between low drifting errors and high compression efficiency. Furthermore, temporal scalability is introduced into the MB-based PFGS, which forms the MB-based PFGST scheme. Two coding modes are proposed for coding the temporal enhancement MBs. Since it would not cause any error propagation if using the high quality reference in the temporal enhancement MB coding, the coding efficiency of the PFGST is highly improved by always choosing the most suitable reference for the temporal scalable coding. Experimental results show that the MB-based PFGST video coding scheme can significantly improve the coding efficiency up to 2.8dB compared with the FGST scheme adopted in MPEG-4, while supporting full SNR, full temporal, and hybrid SNR-temporal scalabilities according to the different requirements from the channels, the clients or the servers. 1. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang |
ICME | 1 |