EDBT 2026 Demo / reviewers in the wild / expert
Yang Cao 0010
dblp:25/7045-10
· DBLP profile ↗
147ranked-venue papers
5as first author
95since 2021 · last 2026
0000-0002-2891-4379ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 3 first-author · 52 since 2021Artificial intelligence and machine learning · 81 · 2 first-author · 63 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E-MaT: Event-oriented Mamba for Egocentric Point TrackingabstractEgocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods struggle in highly dynamic settings—a common challenge in first-person perspectives, where the head-mounted camera undergoes frequent and abrupt rotations, resulting in high angular velocities, motion blur, and large inter-frame displacements. In contrast, event cameras capture motion at microsecond temporal resolution, naturally avoiding blur and delivering low-latency, high-fidelity cues crucial for egocentric point tracking. Moreover, rapid egocentric motion disrupts local smoothness, breaking the assumption that spatially adjacent regions share similar motion. Event dynamics expose global motion trends, guiding coherent modeling and consistent feature flow. Therefore, this paper proposes a mamba-based tracking framework that constructs feature modeling paths aligned with the dominant motion trend extracted from events, and modulates feature propagation along these paths based on local motion intensity, enhancing stability by suppressing unreliable signals and emphasizing consistent cues. Additionally, a motion-adaptive suppression module enhances temporal robustness by adaptively suppressing correlation features based on motion intensity variations, mitigating the effects of intensity fluctuations and partial observability. To facilitate research in this domain, a multimodal dataset named DVS-EgoPoints with both events and videos for egocentric point tracking is collected. Experiments on the DVS-EgoPoints dataset and a simulation benchmark demonstrate superior performance over state-of-the-art methods, especially under challenging motion and occlusion conditions. Wei Zhai, Yang Cao 0010, Bin Li 0025, Zhengjun Zha |
AAAI | 4 |
| 2026 | I²B-LPO: Latent Policy Optimization via Iterative Information BottleneckabstractHuilin Deng, Hongchen Luo, Yue Zhu, Long Li, Zhuoyue Chen, Xinghao Zhao, Ming LI, Chuyang Zhao, Jihai Zhang, MengChang Wang, Yang Cao, Yu Kang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Huilin Deng, Hongchen Luo, Zhuoyue Chen, Xinghao Zhao, Chuyang Zhao, Mengchang Wang, Yang Cao 0010, Yu Kang 0001 |
ACL (1) | 11 |
| 2026 | Visual-Geometric Collaborative Guidance for Affordance Learning
Hongchen Luo, Wei Zhai, Jiao Wang 0002, Yang Cao 0010, Zhengjun Zha |
Int. J. Comput. Vis. | 4 |
| 2026 | Efficient Real-World Image Super-Resolution Via Adaptive Directional Gradient Convolution
Long Peng 0003, Zhanfeng Feng, Renjing Pei, Wenbo Li 0002, Jiaming Guo, Xueyang Fu, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
Int. J. Comput. Vis. | 8 |
| 2026 | Exploring a novel data- and parameter-efficient fine-tuning method for robust face recognition
Yin Lin, Qidong Huang, Yang Cao 0010, Zengfu Wang |
Pattern Recognit. | 7 |
| 2026 | VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) enables the inspection of unseen objects by bridging textual prompts and visual features, showing great potential in flexible manufacturing. While existing ZSAD methods rely on predefined prompts and struggle with unseen defects, Multimodal Large Language Models (MLLMs) offer promising solutions through their generative and interpretative capabilities. However, adapting MLLMs to Industrial Anomaly Detection (IAD) remains challenging due to fine-grained anomaly patterns and subtle visual distinctions. We propose VMAD (Visual-enhanced MLLM Anomaly Detection), a framework that enriches MLLM with visual IAD knowledge through two key components: a Defect-Sensitive Structure Learning scheme that transfers patch-similarities for improved discrimination, and a Locality-enhanced Token Compression that leverages multi-level local features for fine-grained detection. We also introduce RIAD, a comprehensive IAD dataset with detailed anomaly annotations. Extensive experiments on MVTec-AD, Visa, WFDD, and RIAD demonstrate VMAD’s superior performance. The dataset and code will be publicly available at https://github.com/denghuilin-cyber/VMAD. Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Corrections to "VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection"abstractIn the above article [1], an earlier draft of Fig. 6 was inadvertently included. The correct Fig. 6 is presented on the next page.Fig. 6.Zero-shot anomaly segmentation on MVTec-AD, WFDD, and ViSA datasets. Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Modeling Cross-Modal Semantic Transformations From Coarse to Fine in CLIPabstractVision-Language Models (VLMs) like CLIP have advanced image representation through open-vocabulary semantic alignment. Yet, existing few-shot transfer learning methods largely overlook the intrinsic interdependencies between text and image embeddings, limiting their ability to fully transfer CLIP’s pretrained capabilities. To address this gap, we propose Hyperspherical Interpolation Variational Encoding (HIVE), a novel method for few-shot image classification. Our core idea is to shift away from directly training feature extraction capabilities for downstream tasks, and instead focus on exploring the semantic transformation relationships between upstream and downstream tasks. By modeling semantics from coarse to fine granularity, HIVE enables the transfer of original feature extraction and modality alignment capabilities to downstream tasks. Extensive experiments on eight established benchmarks, including CUB and EuroSAT, validate HIVE’s efficacy, achieving up to 46.2% and 80.0% improvements over the original CLIP in 1-shot and 16-shot classification tasks, respectively. Our work underscores the importance of preserving pretrained geometric constraints while exploiting semantic hierarchies for effective few-shot adaptation, providing a principled approach for vision-language model customization. Ziqi Peng, Yang Cao 0010, Yu Kang 0001, Wenjun Lv |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Boosting Image De-Raining via Central-Surrounding Synergistic ConvolutionabstractRainy images suffer from quality degradation due to the synergistic effect of rain streaks and accumulation. The rain streaks are anisotropic and show a specific directional arrangement, while the rain accumulation is isotropic and shows a consistent concentration distribution in local regions. This distribution difference makes unified representation learning for rain streaks and accumulation challenging, which may lead to structure distortion and contrast degradation in the deraining results. To address this problem, a central-surrounding mechanism inspired Synergistic Convolution (SC) is proposed to extract rain streaks and accumulation features simultaneously. Specifically, the SC consists of two parallel novel convolutions: Central-Surrounding Difference Convolution (CSD) and Central-Surrounding Addition Convolution (CSA). In CSD, the difference operation between central and surrounding pixels is injected into the feature extraction process of convolution to perceive the direction distribution of rain streaks. In CSA, the addition operation between central and surrounding pixels is injected into the feature extraction process of convolution to facilitate the modeling of rain accumulation properties. The SC can be used as a general unit to substitute Vanilla Convolution (VC) in current de-raining networks to boost performance. To reduce computational costs, CSA and CSD in SC are merged into a single VC kernel by our parameter equivalent transformation before inferencing. Evaluations of twelve de-raining methods on nine public datasets demonstrate that our proposed SC can comprehensively improve the performance of twelve de-raining networks under various rainy conditions without changing the original network structure or introducing extra computational costs. Even for the current SOTA methods, SC can further achieve SOTA++ performance. The source codes will be publicly available. Long Peng 0003, Yang Wang 0015, Xin Di, Peizhe Xia, Xueyang Fu, Yang Cao 0010, Zhengjun Zha |
AAAI | 6 |
| 2025 | QMambaBSR: Burst Image Super-Resolution with Query State Space ModelabstractBurst super-resolution (BurstSR) aims to reconstruct high-resolution images by fusing subpixel details from multiple low-resolution burst frames. The primary challenge lies in effectively extracting useful information while mitigating the impact of high-frequency noise. Most existing methods rely on frame-by-frame fusion, which often struggles to distinguish informative subpixels from noise, leading to suboptimal performance. To address these limitations, we introduce a novel Query Mamba Burst Super-Resolution (QMambaBSR) network. Specifically, we observe that sub-pixels have consistent spatial distribution while noise appears randomly. Considering the entire burst sequence during fusion allows for more reliable extraction of consistent subpixels and better suppression of noise outliers. Based on this, a Query State Space Model (QSSM) is proposed for both inter-frame querying and intra-frame scanning, enabling a more efficient fusion of useful subpixels. Additionally, to overcome the limitations of static upsampling methods that often result in over-smoothing, we propose an Adaptive Upsampling (AdaUp) module that dynamically adjusts the upsampling kernel to suit the characteristics of different burst scenes, achieving superior detail reconstruction. Extensive experiments on four benchmark datasets—spanning both synthetic and real-world images—demonstrate that QMambaBSR outperforms existing state-of-the-art methods. Xin Di, Long Peng 0003, Peizhe Xia, Wenbo Li 0002, Renjing Pei, Yang Cao 0010, Yang Wang 0015, Zhengjun Zha |
CVPR | 6 |
| 2025 | Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image CaptioningabstractGenerating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this paper, we introduce a detailed caption benchmark, termed as CompreCap, to evaluate the visual context from a directed scene graph view. Concretely, we first manually segment the image into semantically meaningful regions (i.e., semantic segmentation mask) according to common-object vocabulary, while also distinguishing attributes of objects within all those regions. Then directional relation labels of these objects are annotated to compose a directed scene graph that can well encode rich compositional information of the image. Based on our directed scene graph, we develop a pipeline to assess the generated detailed captions from LVLMs on multiple levels, including the object-level coverage, the accuracy of attribute descriptions, the score of key relationships, etc. Experimental results on the CompreCap dataset confirm that our evaluation method aligns closely with human evaluation scores across LVLMs. We have released the code and the dataset here to support the community. Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu 0001, Wei Zhai, Yang Cao 0010, Yujun Shen, Zhengjun Zha |
CVPR | 8 |
| 2025 | GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance GroundingabstractOpen-Vocabulary 3D object affordance grounding aims to anticipate "action possibilities" regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages that depict interactions with 3D geometries to introduce external interaction priors. However, they are still vulnerable to a limited semantic space by failing to leverage implied invariant geometries and potential interaction intentions. Normally, humans address complex tasks through multi-step reasoning and respond to diverse situations by leveraging associative and analogical thinking. In light of this, we propose GREAT (GeometRy-intEntion collAboraTive inference) for Open-Vocabulary 3D Object Affordance Grounding, a novel framework that mines the object invariant geometry attributes and performs analogically reason in potential interaction scenarios to form affordance knowledge, fully combining the knowledge with both geometries and visual contents to ground 3D object affordance. Besides, we introduce the Point Image Affordance Dataset v2 (PIADv2), the largest 3D object affordance dataset at present to support the task. Extensive experiments demonstrate the effectiveness and superiority of GREAT. The code and dataset are available at https://yawen-shao.github.io/GREAT/. Yawen Shao, Wei Zhai, Yuhang Yang 0002, Hongchen Luo, Yang Cao 0010, Zhengjun Zha |
CVPR | 5 |
| 2025 | Improved Video VAE for Latent Video Diffusion ModelabstractVariational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI’s Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial compression, this paper presents two astonishing findings: (1) The initialization from a well-trained image VAE with the same latent dimensions is not an optimal scheme. (2) The adoption of causal reasoning leads to unequal information interactions and unbalanced performance between frames. To alleviate these problems, we propose a keyframe-based temporal compression (KTC) architecture and a group causal convolution (GCConv) module to further improve video VAE (IV-VAE). Specifically, the KTC architecture divides the latent space into two branches, in which one half completely inherits the compression prior of keyframes from a lower-dimension image VAE while the other half involves temporal compression through 3D group causal convolution, reducing temporal-spatial conflicts and accelerating the convergence speed of video VAE. The GC-Conv in the above 3D half uses standard convolution within each frame group to ensure inter-frame equivalence, and employs causal logical padding between groups to maintain flexibility in processing variable frame video. Extensive experiments on five benchmarks demonstrate the SOTA video reconstruction and generation abilities of our IV-VAE. Pingyu Wu, Kai Zhu 0004, Yu Liu 0063, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
CVPR | 6 |
| 2025 | MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic ModelingabstractRecent advancements in multi-modal large language models have propelled the development of joint probabilistic models capable of both image understanding and generation. However, we have identified that recent methods suffer from loss of image information during understanding task, due to either image discretization or diffusion de-noising steps. To address this issue, we propose a novel Multi-Modal Auto-Regressive (MMAR) probabilistic modeling framework. Unlike discretization line of method, MMAR takes in continuous-valued image tokens to avoid information loss in an efficient way. Differing from diffusion-based approaches, we disentangle the diffusion process from auto-regressive backbone model by employing a lightweight diffusion head on top each auto-regressed image patch embedding. In this way, when the model transits from image generation to understanding through text generation, the backbone model’s hidden representation of the image is not limited to the last denoising step. To successfully train our method, we also propose a theoretically proven technique that addresses the numerical stability issue and a training strategy that balances the generation and understanding task goals. Extensive evaluations on 18 image understanding benchmarks show that MMAR significantly outperforms most of the existing joint multi-modal models, surpassing the method that employs pre-trained CLIP vision encoder. Meanwhile, MMAR is able to generate high quality images. We also show that our method is scalable with larger data and model size. Jian Yang 0003, Dacheng Yin, Yizhou Zhou, Fengyun Rao, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
CVPR | 6 |
| 2025 | MATE: Motion-Augmented Temporal Consistency for Event-Based Point Tracking
Wei Zhai, Yang Cao 0010, Bin Li 0025, Zhengjun Zha |
ICCV | 3 |
| 2025 | Emotive: Event-Guided Trajectory Modeling for 3D Motion Estimation
Zengyu Wan, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
ICCV | 3 |
| 2025 | SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assetsabstract3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall into several paradigms: optimization-based and feed-forward (both single-view regression and multi-view generation with reconstruction). However, they are limited by slow speed, low quality, cascade reasoning, and ambiguity in mapping low-dimensional planes to high-dimensional space due to occlusion and invisibility, respectively. Furthermore, existing 3D human assets remain small-scale, insufficient for large-scale training. To address these challenges, we propose a latent space generation paradigm for 3D human digitization, which involves compressing multi-view images into Gaussians via a UV-structured VAE, along with DiT-based conditional generation, we transform the ill-posed low-to-high-dimensional mapping problem into a learnable distribution shift, which also supports end-to-end inference. In addition, we employ the multi-view optimization approach combined with synthetic data to construct the HGS-1M dataset, which contains $1$ million 3D Gaussian assets to support the large-scale training. Experimental results demonstrate that our paradigm, powered by large-scale training, produces high-quality 3D human Gaussians with intricate textures, facial details, and loose clothing deformation. Yuhang Yang 0002, Fengqi Liu, Yixing Lu, Pingyu Wu, Wei Zhai, Ran Yi 0002, Yang Cao 0010, Lizhuang Ma, Zhengjun Zha, Junting Dong |
ICCV | 8 |
| 2025 | HERO: Human Reaction Generation from VideosabstractHuman reaction generation represents a significant research domain for interactive AI, as humans constantly interact with their surroundings. Previous works focus mainly on synthesizing the reactive motion given a human motion sequence. This paradigm limits interaction categories to human-human interactions and ignores emotions that may influence reaction generation. In this work, we propose to generate 3D human reactions from RGB videos, which involves a wider range of interaction categories and naturally provides information about expressions that may reflect the subject's emotions. To cope with this task, we present HERO, a simple yet powerful framework for Human rEaction geneRation from videOs. HERO considers both global and frame-level local representations of the video to extract the interaction intention, and then uses the extracted interaction intention to guide the synthesis of the reaction. Besides, local visual representations are continuously injected into the model to maximize the exploitation of the dynamic properties inherent in videos. Furthermore, the ViMo dataset containing paired Video-Motion data is collected to support the task. In addition to human-human interactions, these video-motion pairs also cover animal-human interactions and scene-human interactions. Extensive experiments demonstrate the superiority of our methodology. The code and dataset will be publicly available at https://jackyu6.github.io/HERO. Chengjun Yu, Wei Zhai, Yuhang Yang 0002, Yang Cao 0010, Zhengjun Zha |
ICCV | 4 |
| 2025 | Towards Realistic Data Generation for Real-World Super-ResolutionabstractExisting image super-resolution (SR) techniques often fail to generalize effectively in complex real-world settings due to the significant divergence between training data and practical scenarios. To address this challenge, previous efforts have either manually simulated intricate physical-based degradations or utilized learning-based techniques, yet these approaches remain inadequate for producing large-scale, realistic, and diverse data simultaneously. In this paper, we introduce a novel Realistic Decoupled Data Generator (RealDGen), an unsupervised learning data generation framework designed for real-world super-resolution. We meticulously develop content and degradation extraction strategies, which are integrated into a novel content-degradation decoupled diffusion model to create realistic low-resolution images from unpaired real LR and HR images. Extensive experiments demonstrate that RealDGen excels in generating large-scale, high-quality paired data that mirrors real-world degradations, significantly advancing the performance of popular SR models on various real-world benchmarks. Long Peng 0003, Wenbo Li 0002, Renjing Pei, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
ICLR | 7 |
| 2025 | FastAno: Accelerating Defect Image Generation with Efficient SamplingabstractDefect inspection faces the challenge of insufficient data. Although existing defect generation methods can produce high-quality defect images, the time-consuming generation process hinders the online availability. To solve it, we propose FastAno, a four-step sampling model for rapid defect generation. Specifically, we first introduce the Adaptive Defect-specific Loss, which calculates region-weighted feature loss to enhance shortcut mapping of defect distribution. Secondly, we propose the Dynamic Attention Optimization Strategy, which enhances the attention activation of the anomaly semantics to improve the generation of defects, while adaptively suppressing the activation of normal semantics to mitigate the degradation of non-defect regions. Extensive experiments on MVTec AD dataset demonstrate that our method achieves significantly faster generation speed while maintaining high generation quality. Haoyu Guan, Qianzi Yu, Kai Zhu 0004, Yang Cao 0010, Yu Kang 0001 |
ICME | 4 |
| 2025 | Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image RestorationabstractImage restoration aims to recover details and enhance contrast in degraded images. With the growing demand for high-quality imaging (e.g., 4K and 8K), achieving a balance between restoration quality and computational efficiency has become increasingly critical. Existing methods, primarily based on CNNs, Transformers, or their hybrid approaches, apply uniform deep representation extraction across the image. However, these methods often struggle to effectively model long-range dependencies and largely overlook the spatial characteristics of image degradation (regions with richer textures tend to suffer more severe damage), making it hard to achieve the best trade-off between restoration quality and efficiency. To address these issues, we propose a novel texture-aware image restoration method, TAMambaIR, which simultaneously perceives image textures and achieves a trade-off between performance and efficiency. Specifically, we introduce a novel Texture-Aware State Space Model, which enhances texture awareness and improves efficiency by modulating the transition matrix of the state-space equation and focusing on regions with complex textures. Additionally, we design a Multi-Directional Perception Block to improve multi-directional receptive fields while maintaining low computational overhead. Extensive experiments on benchmarks for image super-resolution, deraining, and low-light image enhancement demonstrate that TAMambaIR achieves state-of-the-art performance with significantly improved efficiency, establishing it as a robust and efficient framework for image restoration. Long Peng 0003, Xin Di, Zhanfeng Feng, Wenbo Li 0002, Renjing Pei, Yang Wang 0015, Xueyang Fu, Yang Cao 0010, Zhengjun Zha |
IJCAI | 8 |
| 2025 | INFP: INdustrial Video Anomaly Detection via Frequency PrioritizationabstractIndustrial video anomaly detection aims to perform real-time analysis of video streams from industrial production lines and provide anomaly alerts. Conventional video anomaly detection methods focus more on the overall image, as they aim to identify anomalies among multiple normal samples appearing simultaneously. However, industrial scenarios, where the primary focus is on a single type of product, require attention to local areas to capture fine-grained details and specific patterns. Directly applying conventional methods to industrial scenarios can result in an inability to focus on products moving along fixed trajectories, ineffective utilization of their equidistant periodicity, and greater susceptibility to lighting variations. To address these issues, we propose FreqNet, an encoder-decoder framework that learns frequency-domain features from videos to capture periodic and dynamic characteristics, enhancing the model's robustness. Specifically, a trajectory filter is proposed that takes advantage of the significant difference between moving objects and static backgrounds in the frequency domain by assigning higher weights to fixed moving trajectories. Moreover, a multi-feature fusion module is proposed, in which the frequency domain features of the video are first extracted to leverage the unique equidistant periodicity information of videos from industrial production lines. The extracted frequency domain features are subsequently fused with spatio-temporal features and contextual information is further integrated from the fused representation, effectively mitigating the impact of lighting variations on production lines. Extensive experiments on the benchmark IPAD dataset demonstrate the superiority of our proposed method over the state-of-the-art. Qianzi Yu, Kai Zhu 0004, Yang Cao 0010, Yu Kang 0001 |
IJCAI | 3 |
| 2025 | RobustVisH: Robust Visual-Haptic Cross-Modal Recognition under Transmission InterferenceabstractEmbodied AI calls for a reliable, cross-modal object recognition that deeply mines High-Quality (HQ) object appearance (i.e., visual information) and touch details (i.e., haptic information). While in real-world scenarios, cross-modal data is usually degraded due to data acquisition and delivery in complex environments. In this paper, we propose a Robust Visual-Haptic recognition (RobustVisH) model that identifies Low-Quality (LQ) visual-haptic data with transmission distortion for the first time. First, we introduce the WIreless Transmission Interference-based Multi-modal benchmark (WITIM) as a visual-haptic dataset under transmission interference. In particular, the dataset consists of WITIM/AU and WITIM/PHAC-2, in which the original signals are obtained from AU and PHAC-2, respectively. Second, we design a trainable weighted fusion and a Transformer encoder based on the bi-directional self-attention mechanism, enabling RobustVisH to form and learn fused visual-haptic features after modality-specific one-dimensional feature encoding. Third, we employ a covariate shift paradigm, transferring knowledge of RobustVisH from HQ data to LQ data, thereby increasing its robustness against transmission-interference inputs. Experimental results demonstrate that the proposed RobustVisH improves the accuracy of the state-of-the-art method by 2.06% and 9.28% on WITIM/AU and WITIM/PHAC-2, respectively. Source code is available at: https://github.com/lylibylily/RobustVisH. Rouqi Zhang, Chengdi Lu, Hancheng Lu, Yang Cao 0010, Tiesong Zhao |
ACM Multimedia | 4 |
| 2025 | ViewPoint: Panoramic Video Generation with Pretrained Diffusion ModelsabstractPanoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspective data, which constitutes the majority of the training data for modern diffusion models. In this paper, we propose a novel framework utilizing pretrained perspective video models for generating panoramic videos. Specifically, we design a novel panorama representation named ViewPoint map, which possesses global spatial continuity and fine-grained visual details simultaneously. With our proposed Pano-Perspective attention mechanism, the model benefits from pretrained perspective priors and captures the panoramic spatial correlations of the ViewPoint map effectively. Extensive experiments demonstrate that our method can synthesize highly dynamic and spatially consistent panoramic videos, achieving state-of-the-art performance and surpassing previous methods. Zixun Fang, Kai Zhu 0004, Yu Liu 0063, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
NeurIPS | 6 |
| 2025 | PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementabstractMulti-frame video enhancement tasks aim to improve the spatial and temporal resolution and quality of video sequences by leveraging temporal information from multiple frames, which are widely used in streaming video processing, surveillance, and generation. Although numerous Transformer-based enhancement methods have achieved impressive performance, their computational and memory demands hinder deployment on edge devices. Quantization offers a practical solution by reducing the bit-width of weights and activations to improve efficiency. However, directly applying existing quantization methods to video enhancement tasks often leads to significant performance degradation and loss of fine details. This stems from two limitations: (a) inability to allocate varying representational capacity across frames, which results in suboptimal dynamic range adaptation; (b) over-reliance on full-precision teachers, which limits the learning of low-bit student models. To tackle these challenges, we propose a novel quantization method for video enhancement: Progressive Multi-Frame Quantization for Video Enhancement (PMQ-VE). This framework features a coarse-to-fine two-stage process: Backtracking-based Multi-Frame Quantization (BMFQ) and Progressive Multi-Teacher Distillation (PMTD). BMFQ utilizes a percentile-based initialization and iterative search with pruning and backtracking for robust clipping bounds. PMTD employs a progressive distillation strategy with both full-precision and multiple high-bit (INT) teachers to enhance low-bit models' capacity and quality. Extensive experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance across multiple tasks and benchmarks. The code will be made publicly available. Zhanfeng Feng, Long Peng 0003, Xin Di, Wenbo Li 0002, Yulun Zhang 0001, Renjing Pei, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
NeurIPS | 9 |
| 2025 | EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian SplattingabstractScene reconstruction from casually captured videos has wide real-world applications. Despite recent progress, existing methods relying on traditional cameras tend to fail in high-speed scenarios due to insufficient observations and inaccurate pose estimation. Event cameras, inspired by biological vision, record pixel-wise intensity changes asynchronously with high temporal resolution and low latency, providing valuable scene and motion information in blind inter-frame intervals. In this paper, we introduce the event cameras to aid scene construction from a casually captured video for the first time, and propose Event-Aided Free-Trajectory 3DGS, called EF-3DGS, which seamlessly integrates the advantages of event cameras into 3DGS through three key components. First, we leverage the Event Generation Model (EGM) to fuse events and frames, enabling continuous supervision between discrete frames. Second, we extract motion information through Contrast Maximization (CMax) of warped events, which calibrates camera poses and provides gradient-domain constraints for 3DGS. Third, to address the absence of color information in events, we combine photometric bundle adjustment (PBA) with a Fixed-GS training strategy that separates structure and color optimization, effectively ensuring color consistency across different views. We evaluate our method on the public Tanks and Temples benchmark and a newly collected real-world dataset, RealEv-DAVIS. Our method achieves up to 3dB higher PSNR and 40% lower Absolute Trajectory Error (ATE) compared to state-of-the-art methods under challenging high-speed scenarios. Bohao Liao, Wei Zhai, Zengyu Wan, Zhixin Cheng, Wenfei Yang, Yang Cao 0010, Tianzhu Zhang 0001, Zhengjun Zha |
NeurIPS | 6 |
| 2025 | Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized WorldsabstractIn-Context Reinforcement Learning (ICRL) enables agents to learn automatically and on-the-fly from their interactive experiences. However, a major challenge in scaling up ICRL is the lack of scalable task collections. To address this, we propose the procedurally generated tabular Markov Decision Processes, named AnyMDP. Through a carefully designed randomization process, AnyMDP is capable of generating high-quality tasks on a large scale while maintaining relatively low structural biases. To facilitate efficient meta-training at scale, we further introduce decoupled policy distillation and induce prior information in the ICRL framework. Our results demonstrate that, with a sufficiently large scale of AnyMDP tasks, the proposed model can generalize to tasks that were not considered in the training set through versatile in-context learning paradigms. The scalable task set provided by AnyMDP also enables a more thorough empirical investigation of the relationship between data distribution and ICRL performance. We further show that the generalization of ICRL potentially comes at the cost of increased task diversity and longer adaptation periods. This finding carries critical implications for scaling robust ICRL capabilities, highlighting the necessity of diverse and extensive task design, and prioritizing asymptotic performance over few-shot adaptation. Fan Wang 0021, Pengtao Shao, Bo Yu 0014, Shaoshan Liu, Ning Ding 0003, Yang Cao 0010, Yu Kang 0001, Haifeng Wang 0001 |
NeurIPS | 7 |
| 2025 | Transfer learning with a spatiotemporal graph convolution network for city flow predictionabstractRecently, deep learning based city flow prediction has been extensively used in the establishment of smart cities. These methods are data-hungry, making them unscalable to areas lacking data. Although transfer learning can use data-rich source domains to assist target domain cities in city flow prediction, the performance of existing methods cannot meet the needs of actual use, because the long-distance road network connectivity is ignored. To solve this problem, we propose a transfer learning method based on spatiotemporal graph convolution, in which we construct a co-occurrence space between the source and target domains, and then align the mapping of the source and target domains’ data in this space, to achieve the transfer learning of the source city flow prediction model on the target domain. Specifically, a dynamic spatiotemporal graph convolution module along with a temporal encoder is devised to simultaneously capture the concurrent spatiotemporal features, which implies the inherent relationship among the road network structures, human travel habits, and city bike flow. Then, these concurrent features are leveraged as cross-city invariant representations and nonlinearly spanned to a co-occurrence space. The target domain features are thereby aligned with the source domain features in the co-occurrence space by using a Mahalanobis distance loss, to achieve cross-city bike flow prediction. The proposed method is evaluated on the public bike flow datasets in Chicago, New York, and Washington in 2015, and significantly outperforms state-of-the-art techniques. Binkun Liu, Yu Kang 0001, Yang Cao 0010, Yun-Bo Zhao, Zhenyi Xu |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | Multisensor contrast neural network for remaining useful life prediction of rolling bearings under scarce labeled dataabstractPredicting remaining useful life (RUL) of bearings under scarce labeled data is significant for intelligent manufacturing. Current approaches typically encounter the challenge that different degradation stages have similar behaviors in multisensor scenarios. Given that cross-sensor similarity improves the discrimination of degradation features, we propose a multisensor contrast method for RUL prediction under scarce RUL-labeled data, in which we use cross-sensor similarity to mine multisensor similar representations that indicate machine health condition from rich unlabeled sensor data in a co-occurrence space. Specifically, we use ResNet18 to span the features of different sensors into the co-occurrence space. We then obtain multisensor similar representations of abundant unlabeled data through alternate contrast based on cross-sensor similarity in the co-occurrence space. The multisensor similar representations indicate the machine degradation stage. Finally, we focus on finetuning these similar representations to achieve RUL prediction with limited labeled sensor data. The proposed method is evaluated on a publicly available bearing dataset, and the results show that the mean absolute percentage error is reduced by at least 0.058, and the score is improved by at least 0.122 compared with those of state-of-the-art methods. Binkun Liu, Zhenyi Xu, Yu Kang 0001, Yang Cao 0010, Yun-Bo Zhao |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | FSPDD: A double-branch attention guided network for few-shot PCB defect detectionabstractAbstract During the production of printed circuit board (PCB), there will be defects due to inappropriate operations, which will affect the use of electronic products. Majority defect detection methods cost a large number of annotated samples to train detection models. However, PCB defect samples are difficult to collect. Moreover, existing few-shot object detection methods tend to extracting low-level features from support and query images via the shared backbone such as ResNet-50. However, it is not sufficient to obtain fine-grained prior guidance. To address the above issues, we propose a few-shot PCB defect detection model with double-branch attention. Specifically, the joint attention enhancement (JAE) module is proposed to fully mine effective information of query PCB images in multiple dimensions to enhance the representation of latent defects. Then, the multi-scale guidance (MSG) module is proposed to integrate prior knowledge within support PCB images into vectors to reweight query PCB images. Experiments on the PCB defect dataset demonstrate that AP of FSPDD outperforms state-of-the-art methods under different shot settings (k=1,2,3,5,10,30) and our proposed FSPDD has a good generalization ability, in which AP reachs 0.273 when $$k=30$$ k = 30 and is 5.28% higher than SOTA methods. Kehao Shi, Zhenyi Xu, Yang Cao 0010, Lijun Zhao 0003, Yu Kang 0001 |
Multim. Tools Appl. | 3 |
| 2025 | Spatiotemporal Imputation of Traffic Emissions With Self-Supervised Diffusion ModelabstractThe comprehensive regulatory oversight of traffic emissions frequently encounters the missing not-at-random (MNAR) pattern, characterized by the long-term block missing in adjacent road segments, arising from insufficient monitoring points and nonuniform spatiotemporal distribution. The spatiotemporal block missing simultaneously disrupts the spatiotemporal correlation, introducing significant biases in spatiotemporal modeling for incomplete data. The emerging diffusion model recovers the information of the missing regions in a self-supervised manner and focuses on the generation process of the missing regions to address biases. However, the dynamics and spatiotemporal heterogeneity of traffic emissions limit its applicability in unknown spatiotemporal missing. To address this issue, this article proposes a novel progressive Diffusion Model-based framework for SpatioTemporal Imputation of traffic emissions (STI-dm). Specifically, a self-supervised masked training strategy is first devised to construct the nonlocal similarity prior of traffic emission data, explicitly introducing the MNAR missing mechanism for the diffusion process. Furthermore, an enhanced approach of noise injection and supervised denoising is adopted to rectify misconceptions of nonlocal alignment, decreasing modeling biases associated with incomplete data in the generation process. The imputation and prior modeling processes are progressively performed until obtaining stable results, and each of the preceding modeling processes benefits from the gradual improvement results in the other. Experimental evidence indicates that STI-dm surpasses the current state-of-the-art algorithms in scenarios with intricate spatiotemporal patterns and varying rates of missing data. Lihong Pei, Yang Cao 0010, Yu Kang 0001, Zhenyi Xu, Qianming Liu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Tackling Event-Based Lip-Reading by Exploring Multigrained Spatiotemporal CluesabstractAutomatic lip-reading (ALR) is the task of recognizing words based on visual information obtained from the speaker's lip movements. In this study, we introduce event cameras, a novel type of sensing device, for ALR. Event cameras offer both technical and application advantages over conventional cameras for ALR due to their higher temporal resolution, less redundant visual information, and lower power consumption. To recognize words from the event data, we propose a novel multigrained spatiotemporal features learning framework, which is capable of perceiving fine-grained spatiotemporal features from microsecond time-resolved event data. Specifically, we first convert the event data into event frames of multiple temporal resolutions to avoid losing too much visual information at the event representation stage. Then, they are fed into a multibranch subnetwork where the branch operating on low-rate frames can perceive spatially complete but temporally coarse features, while the branch operating on high frame rate can perceive spatially coarse but temporally fine features. Thus, fine-grained spatial and temporal features can be simultaneously learned by integrating the features perceived by different branches. Furthermore, to model the temporal relationships in the event stream, we design a temporal aggregation subnetwork to aggregate the features perceived by the multibranch subnetwork. In addition, we collect two event-based lip-reading datasets (DVS-Lip and DVS-LRW100) for the study of the event-based lip-reading task. Experimental results demonstrate the superiority of the proposed model over the state-of-the-art event-based action recognition models and video-based lip-reading models. Ganchao Tan, Zengyu Wan, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Hypercorrelation Evolution for Video Class-Incremental LearningabstractVideo class-incremental learning aims to recognize new actions while restricting the catastrophic forgetting of old ones, whose representative samples can only be saved in limited memory. Semantically variable subactions are susceptible to class confusion due to data imbalance. While existing methods address the problem by estimating and distilling the spatio-temporal knowledge, we further explores that the refinement of hierarchical correlations is crucial for the alignment of spatio-temporal features. To enhance the adaptability on evolved actions, we proposes a hierarchical aggregation strategy, in which hierarchical matching matrices are combined and jointly optimized to selectively store and retrieve relevant features from previous tasks. Meanwhile, a correlation refinement mechanism is presented to reinforce the bias on informative exemplars according to online hypercorrelation distribution. Experimental results demonstrate the effectiveness of the proposed method on three standard video class-incremental learning benchmarks, outperforming state-of-the-art methods. Code is available at: https://github.com/Lsen991031/HCE Sen Liang, Kai Zhu 0004, Wei Zhai, Yang Cao 0010 |
AAAI | 5 |
| 2024 | LEMON: Learning 3D Human-Object Interaction Relation from 2D ImagesabstractLearning 3D human-object interaction relation is piv-otal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements, e.g., human contact, object affordance, and human-object spatial relation, primarily from the perspective of either the human or the object. Which underexploit certain correlations between the interaction counterparts (human and object), and struggle to address the uncertainty in interactions. Actually, objects' functionalities potentially affect humans' interaction intentions, which reveals what the interaction is. Mean-while, the interacting humans and objects exhibit matching geometric structures, which presents how to interact. In light of this, we propose harnessing these inherent correlations between interaction counterparts to mitigate the uncertainty and jointly anticipate the above interaction el-ements in 3D space. To achieve this, we present LEMON (LEarning 3D huMan-Object iNteraction relation), a unified model that mines interaction intentions of the counter-parts and employs curvatures to guide the extraction of ge-ometric correlations, combining them to anticipate the interaction elements. Besides, the 3D Interaction Relation dataset (3DIR) is collected to serve as the test bed for training and evaluation. Extensive experiments demonstrate the superiority of LEMON over methods estimating each element in isolation. The code and dataset are available at https://yyvhang.github.io/LEMON. Yuhang Yang 0002, Wei Zhai, Hongchen Luo, Yang Cao 0010, Zhengjun Zha |
CVPR | 4 |
| 2024 | Bidirectional Progressive Transformer for Interaction Intention Anticipation
Zichen Zhang 0022, Hongchen Luo, Wei Zhai, Yang Cao 0010, Yu Kang 0001 |
ECCV (59) | 4 |
| 2024 | FC3DNET: A Fully Connected Encoder-Decoder for Efficient DemoiréingabstractMoiré patterns are commonly seen when taking photos of screens. Camera devices usually have limited hardware performance but take high-resolution photos. However, users are sensitive to the photo processing time, which presents a hardly considered challenge of efficiency for demoiréing methods. To balance the network speed and quality of results, we propose a Fully Connected enCoder-deCoder based Demoiréing Network (FC3DNet). FC3DNet utilizes features with multiple scales in each stage of the decoder for comprehensive information, which contains long-range patterns as well as various local moiré styles that both are crucial aspects in demoiréing. Besides, to make full use of multiple features, we design a Multi-Feature Multi-Attention Fusion (MFMAF) module to weigh the importance of each feature and compress them for efficiency. These designs enable our network to achieve performance comparable to state-of-the-art (SOTA) methods in real-world datasets while utilizing only a fraction of parameters, FLOPs, and runtime. Zhibo Du, Long Peng 0003, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
ICIP | 4 |
| 2024 | EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsabstractUnderstanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction specifically manifests in 3D space is also crucial, which links the perception and operation. Existing methods primarily leverage observations of HOI to capture interaction regions from an exocentric view. However, incomplete observations of interacting parties in the egocentric view introduce ambiguity between visual observations and interaction contents, impairing their efficacy. From the egocentric view, humans integrate the visual cortex, cerebellum, and brain to internalize their intentions and interaction concepts of objects, allowing for the pre-formulation of interactions and making behaviors even when interaction regions are out of sight. In light of this, we propose harmonizing the visual appearance, head motion, and 3D object to excavate the object interaction concept and subject intention, jointly inferring 3D human contact and object affordance from egocentric videos. To achieve this, we present EgoChoir, which links object structures with interaction contexts inherent in appearance and head motion to reveal object affordance, further utilizing it to model human contact. Additionally, a gradient modulation is employed to adopt appropriate clues for capturing interaction regions across various egocentric scenarios. Moreover, 3D contact and affordance are annotated for egocentric videos collected from Ego-Exo4D and GIMO to support the task. Extensive experiments on them demonstrate the effectiveness and superiority of EgoChoir. Yuhang Yang 0002, Wei Zhai, Chengfeng Wang, Chengjun Yu, Yang Cao 0010, Zhengjun Zha |
NeurIPS | 5 |
| 2024 | Physics-informed deep Koopman operator for Lagrangian dynamic systems
Yang Cao 0010, Shaofeng Chen, Yu Kang 0001 |
Sci. China Inf. Sci. | 2 |
| 2024 | Grounded Affordance from Exocentric View
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2024 | Background Activation Suppression for Weakly Supervised Object Localization and Semantic Segmentation
Wei Zhai, Pingyu Wu, Kai Zhu 0004, Yang Cao 0010, Feng Wu 0001, Zhengjun Zha |
Int. J. Comput. Vis. | 4 |
| 2024 | A seq2seq learning method for microscopic emission estimation of on-road vehicles
Zhen-Yi Zhao, Yang Cao 0010, Zhenyi Xu, Yu Kang 0001 |
Neural Comput. Appl. | 2 |
| 2024 | On Exploring Multiplicity of Primitives and Attributes for Texture Recognition in the WildabstractTexture recognition is a challenging visual task since its multiple primitives or attributes can be perceived from the texture image under different spatial contexts. Existing approaches predominantly built upon CNN incorporate rich local descriptors with orderless aggregation to capture invariance to the spatial layout. However, these methods ignore the inherent structure relation organized by primitives and the semantic concept described by attributes, which are critical cues for texture representation. In this paper, we propose a novel Multiple Primitives and Attributes Perception network (MPAP) that extracts features by modeling the relation of bottom-up structure and top-down attribute in a multi-branch unified framework. A bottom-up process is first proposed to capture the inherent relation of various primitive structures by leveraging structure dependency and spatial order information. Then, a top-down process is introduced to model the latent relation of multiple attributes by transferring attribute-related features between adjacent branches. Moreover, an augmentation module is devised to bridge the gap between high-level attributes and low-level structure features. MPAP can learn representation through jointing bottom-up and top-down processes in a mutually reinforced manner. Experimental results on six challenging texture datasets demonstrate the superiority of MPAP over state-of-the-art methods in terms of accuracy, robustness, and efficiency. Wei Zhai, Yang Cao 0010, Jing Zhang 0037, Haiyong Xie 0001, Dacheng Tao, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | TF²: Few-Shot Text-Free Training-Free Defect Image Generation for Industrial Anomaly InspectionabstractAnomaly inspection aims at identifying various defects in real time on modern industrial production lines. However, due to insufficient anomaly data, existing detectors cannot effectively accomplish the classification of defects, thereby failing to provide guidance for subsequent production. To address it, we propose TF2, a few-shot text-free training-free defect image generation method, which jointly models the image distribution of class-agnostic defects and backgrounds, achieving efficient semantic enhancement. Firstly, we propose the Response Alignment Strategy, which merges the reversed latent space of both defect-free and defective samples, generating new defect images not limited to textual descriptions yet with consistent content. Moreover, we introduce the Defect Moving Strategy and the Regional Average Loss to merge the reversed latent space of the moving areas and enhance the variability of detail features, increasing both the location and content diversity of defects. Extensive experiments demonstrate the superiority of our model over the state-of-the-art competitors. The metrics indicate that our generated anomaly data focuses on balancing both image quality and diversity, effectively improving the performance of downstream anomaly inspection tasks. Qianzi Yu, Kai Zhu 0004, Yang Cao 0010, Feijie Xia, Yu Kang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Event-Based Optical Flow via Transforming Into Motion-Dependent ViewabstractEvent cameras respond to temporal dynamics, helping to resolve ambiguities in spatio-temporal changes for optical flow estimation. However, the unique spatio-temporal event distribution challenges the feature extraction, and the direct construction of motion representation through the orthogonal view is less than ideal due to the entanglement of appearance and motion. This paper proposes to transform the orthogonal view into a motion-dependent one for enhancing event-based motion representation and presents a Motion View-based Network (MV-Net) for practical optical flow estimation. Specifically, this motion-dependent view transformation is achieved through the Event View Transformation Module, which captures the relationship between the steepest temporal changes and motion direction, incorporating these temporal cues into the view transformation process for feature gathering. This module includes two phases: extracting the temporal evolution clues by central difference operation in the extraction phase and capturing the motion pattern by evolution-guided deformable convolution in the perception phase. Besides, the MV-Net constructs an eccentric downsampling process to avoid response weakening from the sparsity of events in the downsampling stage. The whole network is trained end-to-end in a self-supervised manner, and the evaluations conducted on four challenging datasets reveal the superior performance of the proposed model compared to state-of-the-art (SOTA) methods. Zengyu Wan, Ganchao Tan, Yang Wang 0015, Wei Zhai, Yang Cao 0010, Zhengjun Zha |
IEEE Trans. Image Process. | 5 |
| 2024 | E-MLB: Multilevel Benchmark for Event-Based Camera DenoisingabstractEvent cameras, such as dynamic vision sensors (DVS), are biologically inspired vision sensors that have advanced over conventional cameras in high dynamic range, low latency and low power consumption, showing great application potential in many fields. Event cameras are more sensitive to junction leakage current and photocurrent as they output differential signals, losing the smoothing function of the integral imaging process in the RGB camera. The logarithmic conversion further amplifies noise, especially in low-contrast conditions. Recently, researchers proposed a series of datasets and evaluation metrics but limitations remain: 1) the existing datasets are small in scale and insufficient in noise diversity, which cannot reflect the authentic working environments of event cameras; and 2) the existing denoising evaluation metrics are mostly referenced evaluation metrics, relying on APS information or manual annotation. To address the above issues, we construct a large-scale event denoising dataset (multilevel benchmark for event denoising, E-MLB) for the first time, which consists of 100 scenes, each with four noise levels, that is 12 times larger than the largest existing denoising dataset. We also propose the first nonreference event denoising metric, the event structural ratio (ESR), which measures the structural intensity of given events. ESR is inspired by the contrast metric, but is independent of the number of events and projection direction. Based on the proposed benchmark and ESR, we evaluate the most representative denoising algorithms, including classic and SOTA, and provide denoising baselines under various scenes and noise levels. The corresponding results and codes are available athttps://github.com/KugaMaxx/cuke-emlb. Saizhe Ding, Jinze Chen, Yang Wang 0015, Yu Kang 0001, Yang Cao 0010 |
IEEE Trans. Multim. | 7 |
| 2024 | Lightweight Adaptive Feature De-Drifting for Compressed Image ClassificationabstractJPEG is a widely used compression scheme to efficiently reduce the volume of the transmitted images at the expense of visual perception drop. The artifacts appear among blocks due to the information loss in the compression process, which not only affects the quality of images but also harms the subsequent high-level tasks in terms of feature drifting. High-level vision models trained on high-quality images will suffer performance degradation when dealing with compressed images, especially on mobile devices. In recent years, numerous learning-based JPEG artifacts removal methods have been proposed to handle visual artifacts. However, it is not an ideal choice to use these JPEG artifacts removal methods as a pre-processing for compressed image classification for the following reasons: 1) These methods are designed for human vision rather than high-level vision models. 2) These methods are not efficient enough to serve as a pre-processing on resource-constrained devices. To address these issues, this paper proposes a novel lightweight adaptive feature de-drifting module (AFD-Module) to boost the performance of pre-trained image classification models when facing compressed images. First, a Feature Drifting Estimation Network (FDE-Net) is devised to generate the spatial-wise Feature Drifting Map (FDM) in the DCT domain. Next, the estimated FDM is transmitted to the Feature Enhancement Network (FE-Net) to generate the mapping relationship between degraded features and corresponding high-quality features. Specially, a simple but effective RepConv block equipped with structural re-parameterization is utilized in FE-Net, which enriches feature representation in the training phase while keeping efficiency in the deployment phase. After training on limited compressed images, the AFD-Module can serve as a “plug-and-play” module for pre-trained classification models to improve their performance on compressed images. Experiments on images compressed once (i.e.ImageNet-C) and multiple times demonstrate that our proposed AFD-Module can comprehensively improve the accuracy of the pre-trained classification models and significantly outperform the existing methods. Long Peng 0003, Yang Cao 0010, Yuejin Sun, Yang Wang 0015 |
IEEE Trans. Multim. | 2 |
| 2024 | Learning Visual Affordance Grounding From Demonstration VideosabstractVisual affordance grounding aims to segment all possible interaction regions between people and objects from an image/video, which benefits many applications, such as robot grasping and action recognition. Prevailing methods predominantly depend on the appearance feature of the objects to segment each region of the image, which encounters the following two problems: 1) there are multiple possible regions in an object that people interact with and 2) there are multiple possible human interactions in the same object region. To address these problems, we propose a hand-aided affordance grounding network (HAG-Net) that leverages the aided clues provided by the position and action of the hand in demonstration videos to eliminate the multiple possibilities and better locate the interaction regions in the object. Specifically, HAG-Net adopts a dual-branch structure to process the demonstration video and object image data. For the video branch, we introduce hand-aided attention to enhance the region around the hand in each video frame and then use the long short-term memory (LSTM) network to aggregate the action features. For the object branch, we introduce a semantic enhancement module (SEM) to make the network focus on different parts of the object according to the action classes and utilize a distillation loss to align the output features of the object branch with that of the video branch and transfer the knowledge in the video branch to the object branch. Quantitative and qualitative evaluations on two challenging datasets show that our method has achieved state-of-the-art results for affordance grounding. The source code is available at: https://github.com/lhc1224/HAG-Net. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Self-Supervised Spatiotemporal Clustering of Vehicle Emissions With Graph Convolutional NetworkabstractSpatiotemporal clustering of vehicle emissions, which reveals the evolution pattern of air pollution from road traffic, is a challenging representation learning task due to the lack of supervision. Some recent work building upon graph convolutional network (GCN) models the intrinsic spatiotemporal correlations among the nodes in road networks as graph representations for clustering. However, these existing methods ignore the interactions between spatial and temporal variations in vehicle emissions, resulting in incomplete descriptions and inaccurate detection of the evolution pattern of air pollution. To address this issue, this article proposes a two-way self-supervised spatiotemporal representation learning scheme, in which the temporal and spatial features are progressively learned in a mutually reinforced manner. Our proposed method is based on the observation that though the variation in vehicle emissions in the road network is consistent in the spatial and temporal domains, its expression is more distinct in temporal sequences. To this end, the input emission data are first projected into an initial temporal representation space spanned by the captured features from a pretrained BiLSTM network. Then the generated distribution of temporal features is used to construct an objective constraint for high-purity clustering through a two-way self-supervised mechanism, which is leveraged as a constraint for the feature clustering of a GCN. Furthermore, to eliminate the initial errors, a joint optimization scheme is presented to generate the decoupled clustering results through the progressive refinement of representation and clustering. Our proposed method is evaluated on the traffic emission dataset of Xian city in 2020, and the experimental results have demonstrated the superiority against the state-of-the-art. Lihong Pei, Yang Cao 0010, Yu Kang 0001, Zhenyi Xu, Zhen-Yi Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Uncertainty-Aware Optimal Transport for Semantically Coherent Out-of-Distribution DetectionabstractSemantically coherent out-of-distribution (SCOOD) detection aims to discern outliers from the intended data distribution with access to unlabeled extra set. The coexistence of in-distribution and out-of-distribution samples will exacerbate the model overfitting when no distinction is made. To address this problem, we propose a novel uncertainty-aware optimal transport scheme. Our scheme consists of an energy-based transport (ET) mechanism that estimates the fluctuating cost of uncertainty to promote the assignment of semantic-agnostic representation, and an inter-cluster extension strategy that enhances the discrimination of semantic property among different clusters by widening the corresponding margin distance. Furthermore, a T-energy score is presented to mitigate the magnitude gap between the parallel transport and classifier branches. Extensive experiments on two standard SCOOD benchmarks demonstrate the above-par OOD detection performance, outperforming the state-of-the-art methods by a margin of 27.69% and 34.4% on FPR@95, respectively. Code is available at https://github.com/LuFan3/IET-OOD. Kai Zhu 0004, Wei Zhai, Kecheng Zheng, Yang Cao 0010 |
CVPR | 5 |
| 2023 | Leverage Interactive Affinity for Affordance LearningabstractPerceiving potential “action possibilities” (i.e., affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assignment paradigm and presume that there is a unique relationship between functional region and affordance label, yielding poor performance when adapting to unseen environments with large appearance variations. In this paper, we propose to leverage interactive affinity for affordance learning, i.e. extracting interactive affinity from human-object interaction and transferring it to non-interactive objects. Interactive affinity, which represents the contacts between different parts of the human body and local regions of the target object, can provide inherent cues of interconnectivity between humans and objects, thereby reducing the ambiguity of the perceived action possibilities. Specifically, we propose a pose-aided interactive affinity learning framework that exploits human pose to guide the network to learn the interactive affinity from human-object interactions. Particularly, a keypoint heuristic perception (KHP) scheme is devised to exploit the keypoint association of human pose to alleviate the uncertainties due to interaction diversities and contact occlusions. Besides, a contact-driven affordance learning (CAL) dataset is constructed by collecting and labeling over 5, 000 images. Experimental results demonstrate that our method outperforms the representative models regarding objective metrics and visual quality. Code and dataset: github.com/lhc1224/PIAL-Net. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
CVPR | 4 |
| 2023 | Decoupling-and-Aggregating for Image Exposure CorrectionabstractThe images captured under improper exposure conditions often suffer from contrast degradation and detail distortion. Contrast degradation will destroy the statistical properties of low-frequency components, while detail distortion will disturb the structural properties of high-frequency components, leading to the low-frequency and high-frequency components being mixed and inseparable. This will limit the statistical and structural modeling capacity for exposure correction. To address this issue, this paper proposes to decouple the contrast enhancement and detail restoration within each convolution process. It is based on the observation that, in the local regions covered by convolution kernels, the feature response of low-/high-frequency can be decoupled by addition/difference operation. To this end, we inject the addition/difference operation into the convolution process and devise a Contrast Aware (CA) unit and a Detail Aware (DA) unit to facilitate the statistical and structural regularities modeling. The proposed CA and DA can be plugged into existing CNN-based exposure correction networks to substitute the Traditional Convolution (TConv) to improve the performance. Furthermore, to maintain the computational costs of the network without changing, we aggregate two units into a single TConv kernel using structural re-parameterization. Evaluations of nine methods and five benchmark datasets demonstrate that our proposed method can comprehensively improve the performance of existing methods without introducing extra computational costs compared with the original networks. The codes will be publicly available. Yang Wang 0015, Long Peng 0003, Liang Li 0003, Yang Cao 0010, Zhengjun Zha |
CVPR | 4 |
| 2023 | Self-Organizing Pathway Expansion for Non-Exemplar Class-Incremental LearningabstractNon-exemplar class-incremental learning aims to recognize both the old and new classes without access to old class samples. The conflict between old and new class optimization is exacerbated since the shared neural pathways can only be differentiated by the incremental samples. To address this problem, we propose a novel self-organizing pathway expansion scheme. Our scheme consists of a class-specific pathway organization strategy that reduces the coupling of optimization pathway among different classes to enhance the independence of the feature representation, and a pathway-guided feature optimization mechanism to mitigate the update interference between the old and new classes. Extensive experiments on four datasets demonstrate significant performance gains, outperforming the state-of-the-art methods by a margin of 1%, 3%, 2% and 2%, respectively. Kai Zhu 0004, Kecheng Zheng, Ruili Feng, Deli Zhao, Yang Cao 0010, Zhengjun Zha |
ICCV | 5 |
| 2023 | Spatial-Aware Token for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging task aiming to localize objects with only image-level supervision. Recent works apply visual transformer to WSOL and achieve significant success by exploiting the long-range feature dependency in self-attention mechanism. However, existing transformer-based methods synthesize the classification feature maps as the localization map, which leads to optimization conflicts between classification and localization tasks. To address this problem, we propose to learn a task-specific spatial-aware token (SAT) to condition localization in a weakly supervised manner. Specifically, a spatial token is first introduced in the input space to aggregate representations for localization task. Then a spatial aware attention module is constructed, which allows spatial token to generate foreground probabilities of different patches by querying and to extract localization knowledge from the classification task. Besides, for the problem of sparse and unbalanced pixel-level supervision obtained from the image-level label, two spatial constraints, including batch area loss and normalization loss, are designed to compensate and enhance this supervision. Experiments show that the proposed SAT achieves state-of-the-art performance on both CUB-200 and ImageNet, with 98.45% and 73.13% GT-known Loc, respectively. Even under the extreme setting of using only 1 image per class from ImageNet for training, SAT already exceeds the SOTA method by 2.1% GT-known Loc. Code and models are available at https://github.com/wpy1999/SAT. Pingyu Wu, Wei Zhai, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha |
ICCV | 3 |
| 2023 | Grounding 3D Object Affordance from 2D Interactions in ImagesabstractGrounding 3D object affordance seeks to locate objects’ "action possibilities" regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual affordances with geometry structures, e.g., relying on annotations to declare interactive regions of interest on the object and establishing a mapping between the regions and affordances. However, the essence of learning object affordance is to understand how to use it, and the manner that detaches interactions is limited in generalization. Normally, humans possess the ability to perceive object affordances in the physical world through demonstration images or videos. Motivated by this, we introduce a novel task setting: grounding 3D object affordance from 2D interactions in images, which faces the challenge of anticipating affordance through interactions of different sources. To address this problem, we devise a novel Interaction-driven 3D Affordance Grounding Network (IAG), which aligns the region feature of objects from different sources and models the interactive contexts for 3D object affordance grounding. Besides, we collect a Point-Image Affordance Dataset (PIAD) to support the proposed task. Comprehensive experiments on PIAD demonstrate the reliability of the proposed task and the superiority of our method. The project is available at https://github.com/yyvhang/IAGNet. Yuhang Yang 0002, Wei Zhai, Hongchen Luo, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha |
ICCV | 4 |
| 2023 | Cones: Concept Neurons in Diffusion Models for Customized GenerationabstractHuman brains respond to semantic features of presented stimuli with different neurons. This raises the question of whether deep neural networks admit a similar behavior pattern. To investigate this phenomenon, this paper identifies a small cluster of neurons associated with a specific subject in a diffusion model. We call those neurons the concept neurons. They can be identified by statistics of network gradients to a stimulation connected with the given subject. The concept neurons demonstrate magnetic properties in interpreting and manipulating generation results. Shutting them can directly yield the related subject contextualized in different scenes. Concatenating multiple clusters of concept neurons can vividly generate all related concepts in a single image. Our method attains impressive performance for multi-subject customization, even four or more subjects. For large-scale applications, the concept neurons are environmentally friendly as we only need to store a sparse cluster of int index instead of dense float32 parameter values, reducing storage consumption by 90% compared with previous customized generation methods. Extensive qualitative and quantitative studies on diverse scenarios show the superiority of our method in interpreting and manipulating diffusion models. Ruili Feng, Kai Zhu 0004, Kecheng Zheng, Yu Liu 0063, Deli Zhao, Jingren Zhou 0001, Yang Cao 0010 |
ICML | 9 |
| 2023 | Fusion-Based Low-Light Image Enhancement
Haodian Wang, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
MMM (1) | 3 |
| 2023 | Customizable Image Synthesis with Multiple SubjectsabstractSynthesizing images with user-specified subjects has received growing attention due to its practical applications. Despite the recent success in single subject customization, existing algorithms suffer from high training cost and low success rate along with increased number of subjects. Towards controllable image synthesis with multiple subjects as the constraints, this work studies how to efficiently represent a particular subject as well as how to appropriately compose different subjects. We find that the text embedding regarding the subject token already serves as a simple yet effective representation that supports arbitrary combinations without any model tuning. Through learning a residual on top of the base embedding, we manage to robustly shift the raw subject to the customized subject given various text conditions. We then propose to employ layout, a very abstract and easy-to-obtain prior, as the spatial guidance for subject arrangement. By rectifying the activations in the cross-attention map, the layout appoints and separates the location of different subjects in the image, significantly alleviating the interference across them. Using cross-attention map as the intermediary, we could strengthen the signal of target subjects and weaken the signal of irrelevant subjects within a certain region, significantly alleviating the interference across subjects. Both qualitative and quantitative experimental results demonstrate our superiority over state-of-the-art alternatives under a variety of settings for multi-subject customization. Yujun Shen, Kecheng Zheng, Kai Zhu 0004, Ruili Feng, Yu Liu 0063, Deli Zhao, Jingren Zhou 0001, Yang Cao 0010 |
NeurIPS | 10 |
| 2023 | P3DC-shot: Prior-driven discrete data calibration for nearest-neighbor few-shot classification
Shuangmei Wang, Rui Ma 0011, Tieru Wu, Yang Cao 0010 |
Image Vis. Comput. | 4 |
| 2023 | High-emitter identification for heavy-duty vehicles by temporal optimization LSTM and an adaptive dynamic thresholdabstractHeavy-duty diesel vehicles are important sources of urban nitrogen oxides (NOx) in actual applications for environmental compliance, emitting more than 80% of NOx and more than 90% of particulate matter (PM) in total vehicle emissions. The detection and control of heavy-duty diesel emissions are critical for protecting public health. Currently, vehicles on the road must be regularly tested, every six months or once a year, to filter out high-emission mobile sources at vehicle inspection stations. However, it is difficult to effectively screen high-emission vehicles in time with a long interval between annual inspections, and the fixed threshold cannot adapt to the dynamic changes of vehicle driving conditions. An on-board diagnostic device (OBD) is installed inside the vehicle and can record the vehicle’s emission data in real time. In this paper, we propose a temporal optimization long short-term memory (LSTM) and adaptive dynamic threshold approach to identify heavy-duty high-emitters by using OBD data, which can continuously track and record the emission status in real time. First, a temporal optimization LSTM emission prediction model is established to solve the attention bias discrepancy problem on time steps that is caused by the large number of OBD data streams in practice. Then, the concentration prediction error sequence is detected and distinguished from the anomalous emission contexts using flexible criteria, calculated by an adaptive dynamic threshold with changing driving conditions. Finally, a similarity metric strategy for the time series is introduced to correct some pseudo anomalous results. Experiments on three real OBD time-series emission datasets demonstrate that our method can achieve high accuracy anomalous emission identification. Zhenyi Xu, Renjun Wang, Yang Cao 0010, Yu Kang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | Traffic emission estimation under incomplete information with spatiotemporal convolutional GAN
Zhen-Yi Zhao, Yang Cao 0010, Zhenyi Xu, Yu Kang 0001 |
Neural Comput. Appl. | 2 |
| 2023 | Location-Free Camouflage Generation NetworkabstractCamouflage is a common visual phenomenon, which refers to hiding the foreground objects into the background images, making them briefly invisible to the human eye. Previous work has typically been implemented by an iterative optimization process. However, these methods struggle in 1) efficiently generating camouflage images using foreground and background with flexible structure; 2) camouflaging foreground objects to regions with multiple appearances (e.g.the junction of the vegetation and the mountains), which limit their practical application. To address these problems, this paper proposes a novelLocation-freeCamouflageGenerationNetwork(LCG-Net) that fuse high-level features of foreground and background image, and generate result by one inference. Specifically, aPosition-aligned Structure Fusion(PSF) module is devised to guide structure feature fusion based on the point-to-point structure similarity of foreground and background, and introduce local appearance features point-by-point. To retain the necessary identifiable features, a new immerse loss is adopted under our pipeline, while a background patch appearance loss is utilized to ensure that the hidden objects look continuous and natural at regions with multiple appearances. Experiments show that our method has results as satisfactory as state-of-the-art in the single-appearance regions and are less likely to be completely invisible, but far exceed the quality of the state-of-the-art in the multi-appearance regions. Moreover, our method is hundreds of times faster than previous methods. Benefitting from the unique advantages of our method, we provide some downstream applications for camouflage generation, which show its potential. The related code and dataset will be released athttps://github.com/Tale17/LCG-Net. Wei Zhai, Yang Cao 0010, Zhengjun Zha |
IEEE Trans. Multim. | 3 |
| 2023 | Deep Texton-Coherence Network for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging visual task since the appearance and morphology of foreground objects and background regions are highly similar in nature. Recent CNN-based studies gradually integrated the high-level semantic information and the low-level local features of images through hierarchical and progressive structures to achieve camouflaged object detection. However, these methods ignore thespatial statistical propertiesof the local context, which is a critical cue for distinguishing and describing camouflaged objects. To address this problem, we propose a novel Deep Texton-Coherence Network (DTC-Net) that leverages the spatial organization of textons in the foreground and background regions as discriminative cues for camouflaged object detection. Specifically, a Local Bilinear module (LB) is devised to obtain the robust representation of texton to trivial details and illumination changes, by replacing the classic first-order linearization operations with bilinear second-order statistical operations in the convolution process. Next, these texton representations are associated with a Spatial Coherence Organization module (SCO) to capture irregular spatial coherence via a deformable convolutional strategy, and then the descriptions of the textons extracted by the LB module are used as weights to suppress features that are spatially adjacent but have different representations. Finally, the texton-coherence representation is integrated with the original features at different levels to achieve camouflaged object detection. Evaluation on the three most challenging camouflaged object detection datasets demonstrats the superiority of the proposed model when compared to the state-of-the-art methods. Furthermore, our ablation studies and performance analyses demonstrate the effectiveness of the texton-coherence module. Wei Zhai, Yang Cao 0010, Haiyong Xie 0001, Zhengjun Zha |
IEEE Trans. Multim. | 2 |
| 2023 | Convex Temporal Convolutional Network-Based Distributed Cooperative Learning Control for Multiagent SystemsabstractDue to its great efficiency, scalability, and inclusivity, distributed cooperative learning control has gotten a lot of attention. For complex uncertain multiagent systems, it is challenging to model the uncertainties and exploit the cooperative learning ability of the systems. To address these issues, we proposed a novel convex temporal convolutional network-based distributed cooperative learning control for uncertain discrete-time nonlinear multiagent systems. A new concept of using a convex temporal convolutional network (CTCNet) is proposed for estimating the uncertain agent dynamics in a cooperative way. Unlike previous methods that require adjustment of network weights for different control tasks, the proposed CTCNet can map the high-dimensional input-output space into a deep space spanned by basis features that represent the inherent properties of the system, so it has good robustness for different tasks. Consequently, to improve the control performance, a CTCNet-based distributed cooperative learning control method that shares learned knowledge through the communication topology among adaptive laws of CTCNet is proposed. Furthermore, the asymptotic convergence of system tracking errors to an arbitrarily small neighborhood of the origin is strictly proved. Finally, the simulation results are given to illustrate that our suggested method has higher control accuracy, stronger robustness, and anti-interference ability than the existing methods. Shaofeng Chen, Yu Kang 0001, Jian Di, Pengfei Li 0006, Yang Cao 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | ProgressiveMotionSeg: Mutually Reinforced Framework for Event-Based Motion SegmentationabstractDynamic Vision Sensor (DVS) can asynchronously output the events reflecting apparent motion of objects with microsecond resolution, and shows great application potential in monitoring and other fields. However, the output event stream of existing DVS inevitably contains background activity noise (BA noise) due to dark current and junction leakage current, which will affect the temporal correlation of objects, resulting in deteriorated motion estimation performance. Particularly, the existing filter-based denoising methods cannot be directly applied to suppress the noise in event stream, since there is no spatial correlation. To address this issue, this paper presents a novel progressive framework, in which a Motion Estimation (ME) module and an Event Denoising (ED) module are jointly optimized in a mutually reinforced manner. Specifically, based on the maximum sharpness criterion, ME module divides the input event into several segments by adaptive clustering in a motion compensating warp field, and captures the temporal correlation of event stream according to the clustered motion parameters. Taking temporal correlation as guidance, ED module calculates the confidence that each event belongs to real activity events, and transmits it to ME module to update energy function of motion segmentation for noise suppression. The two steps are iteratively updated until stable motion segmentation results are obtained. Extensive experimental results on both synthetic and real datasets demonstrate the superiority of our proposed approaches against the State-Of-The-Art (SOTA) methods. Jinze Chen, Yang Wang 0015, Yang Cao 0010, Feng Wu 0001, Zhengjun Zha |
AAAI | 3 |
| 2022 | Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental LearningabstractNon-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a novel self-sustaining representation expansion scheme. Our scheme consists of a structure reorganization strategy that fuses main-branch expansion and side-branch updating to maintain the old features, and a main-branch distillation scheme to transfer the invariant knowledge. Furthermore, a prototype selection mechanism is proposed to enhance the discrimination between the old and new classes by selectively incorporating new samples into the distillation process. Extensive experiments on three benchmarks demonstrate significant incremental performance, outperforming the state-of-the-art methods by a margin of 3%, 3% and 6%, respectively. Kai Zhu 0004, Wei Zhai, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha |
CVPR | 3 |
| 2022 | Learning Affordance Grounding from Exocentric ImagesabstractAffordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability that transform the various exocentric interactions to invariant egocentric affordance so as to counter the impact of interactive diversity. To empower an agent with such ability, this paper proposes a task of affordance grounding from exocentric view, i.e., given exocentric human-object interaction and egocentric object images, learning the affordance knowledge of the object and transferring it to the egocentric image using only the affordance label as supervision. To this end, we devise a cross-view knowledge transfer framework that extracts affordance-specific features from exocentric interactions and enhances the perception of affordance regions by preserving affordance correlation. Specifically, an Affordance Invariance Mining module is devised to extract specific clues by minimizing the intra-class differences originated from interaction habits in exocentric images. Besides, an Affordance Co-relation Preserving strategy is presented to perceive and localize affordance by aligning the co-relation matrix of predicted results between the two views. Particularly, an affordance grounding dataset named AGD20K is constructed by collecting and labeling over 20K images from 36 affordance categories. Experimental results demonstrate that our method outperforms the representative models in terms of objective metrics and visual quality. Code: github.com/lhc1224/Cross-View-AG. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
CVPR | 4 |
| 2022 | Multi-grained Spatio-Temporal Features Perceived Network for Event-based Lip-ReadingabstractAutomatic lip-reading (ALR) aims to recognize words using visual information from the speaker's lip movements. In this work, we introduce a novel type of sensing device, event cameras, for the task of ALR. Event cameras have both technical and application advantages over conventional cameras for the ALR task because they have higher temporal resolution, less redundant visual information, and lower power consumption. To recognize words from the event data, we propose a novel Multi-grained Spatio-Temporal Features Perceived Network (MSTP) to perceive fine-grained spatio-temporal features from microsecond time-resolved event data. Specifically, a multi-branch network architecture is designed, in which different grained spatio-temporal features are learned by operating at different frame rates. The branch operating on the low frame rate can perceive spatial complete but temporal coarse features. While the branch operating on the high frame rate can perceive spatial coarse but temporal refinement features. And a message flow module is devised to integrate the features from different branches, leading to perceiving more discriminative spatio-temporal features. In addition, we present the first event-based lip-reading dataset (DVS-Lip) captured by the event camera. Experimental results demonstrated the superiority of the proposed model compared to the state-of-the-art event-based action recognition models and video-based lip-reading models. Ganchao Tan, Yang Wang 0015, Yang Cao 0010, Feng Wu 0001, Zhengjun Zha |
CVPR | 4 |
| 2022 | Background Activation Suppression for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects using only image-level labels. Recently a new paradigm has emerged by generating a foreground prediction map (FPM) to achieve localization task. Existing FPM-based methods use cross-entropy (CE) to evaluate the foreground prediction map and to guide the learning of generator. We argue for using activation value to achieve more efficient learning. It is based on the experimental observation that, for a trained network, CE converges to zero when the foreground mask covers only part of the object region. While activation value increases until the mask expands to the object boundary, which indicates that more object areas can be learned by using activation value. In this paper, we propose a Background Activation Suppression (BAS) method. Specifically, an Activation Map Constraint module (AMC) is designed to facilitate the learning of generator by suppressing the background activation value. Meanwhile, by using the foreground region guidance and the area constraint, BAS can learn the whole region of the object. In the inference phase, we consider the prediction maps of different categories together to obtain the final localization results. Extensive experiments show that BAS achieves significant and consistent improvement over the baseline methods on the CUB-200-2011 and ILSVRC datasets. Code and models are available at github.com/wpy1999IBAS. Pingyu Wu, Wei Zhai, Yang Cao 0010 |
CVPR | 3 |
| 2022 | Dreaming to Prune Image Deraining NetworksabstractConvolutional image deraining networks have achieved great success while suffering from tremendous computational and memory costs. Most model compression methods require original data for iterative fine-tuning, which is limited in real-world applications due to storage, privacy, and transmission constraints. We note that it is overstretched to fine-tune the compressed model using self-collected data, as it exhibits poor generalization over images with different degradation characteristics. To address this problem, we propose a novel data-free compression framework for de-raining networks. It is based on our observation that deep degradation representations can be clustered by degradation characteristics (types of rain) while independent of image content. Therefore, in our framework, we “dream” diverse in-distribution degraded images using a deep inversion paradigm, thus leveraging them to distill the pruned model. Specifically, we preserve the performance of the pruned model in a dual-branch way. In one branch, we invert the pre-trained model (teacher) to reconstruct the degraded inputs that resemble the original distribution and employ the orthogonal regularization for deep features to yield degradation diversity. In the other branch, the pruned model (student) is distilled to fit the teacher's original statistical modeling on these dreamed inputs. Further, an adaptive pruning scheme is proposed to determine the hierarchical sparsity, which alleviates the regression drift of the initial pruned model. Experiments on various deraining datasets demonstrate that our method can reduce about 40% FLOPs of the state-of-the-art models while maintaining comparable performance without original data. Weiqi Zou, Yang Wang 0015, Xueyang Fu, Yang Cao 0010 |
CVPR | 4 |
| 2022 | S2N: Suppression-Strengthen Network for Event-Based Recognition Under Variant Illuminations
Zengyu Wan, Yang Wang 0015, Ganchao Tan, Yang Cao 0010, Zhengjun Zha |
ECCV (3) | 4 |
| 2022 | Towards Data-Efficient Detection Transformers
Wen Wang 0009, Jing Zhang 0037, Yang Cao 0010, Yongliang Shen 0001, Dacheng Tao |
ECCV (9) | 3 |
| 2022 | Fast Adaptive Self-Supervised Underwater Image EnhancementabstractWavelength-dependent light absorption and scattering result in the color cast and contrast degradation, which degrades the visibility of underwater images. Deep-learning-based under-water image enhancement has made significant progress in recent years. However, existing algorithms rely heavily on massive data for training, and the difficulty of collecting massive data in real-world environments limits their applicability. To alleviate this problem, we propose a fast adaptive self-supervised underwater image enhancement method, which models the entire water domain in terms of color and contrast distribution with only a few images. Specifically, to learn the mapping of color correction and contrast enhancement, our framework contains two sub-networks: a Color Mapping Net (CLM-Net) and a Contrast Mapping Net (CTM-Net). The CLM-Net learns the color correction mapping under the constraint of gray world assumption in a self-supervised manner. Besides, the CTM-Net estimates the contrast enhancement mapping under the constraint of the dark channel prior. Thanks to adaptive self-supervised training, we can primarily alleviate the need for collecting massive data and achieve effective underwater image enhancement with only ten images. Experiments demonstrate that the proposed algorithm achieves state-of-the-art performance on four benchmark datasets. Mengxiao Huang, Yang Wang 0015, Weiqi Zou, Yang Cao 0010 |
ICIP | 4 |
| 2022 | FP-DETR: Detection Transformer Advanced by Fully Pre-training
Wen Wang 0009, Yang Cao 0010, Jing Zhang 0037, Dacheng Tao |
ICLR | 2 |
| 2022 | A Two-Layers Super-Resolution Based Generation Adversarial Spatiotemporal Fusion ModelabstractRemote sensing image spatiotemporal fusion (STF) algorism plays an important role by supplementing the lack of original high-resolution remote sensing satellite images in the study scenarios of dense time-series data. In recent years, the deep-learning-based STF algorithm has become a research hotspot with comparatively higher accuracy and robustness. However, due to the lack of sufficient high-quality images for training and the huge resolution gap between low-resolution images and high-resolution images, it is difficult to recover detailed information, especially for areas of land-cover change. In this paper, we propose a two-layers super-resolution based generation adversarial spatiotemporal fusion model(TLSRSTF) using smaller inputs to reduce pressure on data requirements and a mutual affine convolution to reduce model parameters. Specifically, we only use a pair of high-resolution and low-resolution images and a high-resolution image at any time. A spatial degradation consistency is constructed to adaptively determine the ratio of two layers of the super-resolution STF model. The quantitative and qualitative experimental results on public spatiotemporal fusion datasets demonstrate our superiority over the state-of-the-art methods. Shuai Fang, Yang Cao 0010, Jing Zhang 0037 |
IGARSS | 3 |
| 2022 | Long-Range Feature Dependencies Capturing for Low-Resolution Image Classification
Sheng Kang, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
MMM (2) | 3 |
| 2022 | Lightweight Wavelet-Based Network for JPEG Artifacts Removal
Yuejin Sun, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
MMM (2) | 3 |
| 2022 | AS-Net: Class-Aware Assistance and Suppression Network for Few-Shot Learning
Ruijing Zhao, Kai Zhu 0004, Yang Cao 0010, Zhengjun Zha |
MMM (2) | 3 |
| 2022 | Uncertainty-Aware Hierarchical Refinement for Incremental Implicitly-Refined ClassificationabstractIncremental implicitly-refined classification task aims at assigning hierarchical labels to each sample encountered at different phases. Existing methods tend to fail in generating hierarchy-invariant descriptors when the novel classes are inherited from the old ones. To address the issue, this paper, which explores the inheritance relations in the process of multi-level semantic increment, proposes an Uncertainty-Aware Hierarchical Refinement (UAHR) scheme. Specifically, our proposed scheme consists of a global representation extension strategy that enhances the discrimination of incremental representation by widening the corresponding margin distance, and a hierarchical distribution alignment strategy that refines the distillation process by explicitly determining the inheritance relationship of the incremental class. Particularly, the shifting subclasses are corrected under the guidance of hierarchical uncertainty, ensuring the consistency of the homogeneous features. Extensive experiments on widely used benchmarks (i.e., IIRC-CIFAR, IIRC-ImageNet-lite, IIRC-ImageNet-Subset, and IIRC-ImageNet-full) demonstrate the superiority of our proposed method over the state-of-the-art approaches. Jian Yang 0003, Kai Zhu 0004, Kecheng Zheng, Yang Cao 0010 |
NeurIPS | 4 |
| 2022 | Exploring Figure-Ground Assignment Mechanism in Perceptual OrganizationabstractPerceptual organization is a challenging visual task that aims to perceive and group the individual visual element so that it is easy to understand the meaning of the scene as a whole. Most recent methods building upon advanced Convolutional Neural Network (CNN) come from learning discriminative representation and modeling context hierarchically. However, when the visual appearance difference between foreground and background is obscure, the performance of existing methods degrades significantly due to the visual ambiguity in the discrimination process. In this paper, we argue that the figure-ground assignment mechanism, which conforms to human vision cognitive theory, can be explored to empower CNN to achieve a robust perceptual organization despite visual ambiguity. Specifically, we present a novel Figure-Ground-Aided (FGA) module to learn the configural statistics of the visual scene and leverage it for the reduction of visual ambiguity. Particularly, we demonstrate the benefit of using stronger supervisory signals by teaching (FGA) module to perceive configural cues, \ie, convexity and lower region, that human deem important for the perceptual organization. Furthermore, an Interactive Enhancement Module (IEM) is devised to leverage such configural priors to assist representation learning, thereby achieving robust perception organization with complex visual ambiguities. In addition, a well-founded visual segregation test is designed to validate the capability of the proposed FGA mechanism explicitly. Comprehensive evaluation results demonstrate our proposed FGA mechanism can effectively enhance the capability of perception organization on various baseline models. Nevertheless, the model augmented via our proposed FGA mechanism also outperforms state-of-the-art approaches on four challenging real-world applications. Wei Zhai, Yang Cao 0010, Jing Zhang 0037, Zhengjun Zha |
NeurIPS | 2 |
| 2022 | One-Shot Object Affordance Detection in the Wild
Wei Zhai, Hongchen Luo, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2022 | Active Domain Adaptation With Application to Intelligent Logging Lithology IdentificationabstractLithology identification plays an essential role in formation characterization and reservoir exploration. As an emerging technology, intelligent logging lithology identification has received great attention recently, which aims to infer the lithology type through the well-logging curves using machine-learning methods. However, the model trained on the interpreted logging data is not effective in predicting new exploration well due to the data distribution discrepancy. In this article, we aim to train a lithology identification model for the target well using a large amount of source-labeled logging data and a small amount of target-labeled data. The challenges of this task lie in three aspects: 1) the distribution misalignment; 2) the data divergence; and 3) the cost limitation. To solve these challenges, we propose a novel active adaptation for logging lithology identification (AALLI) framework that combines active learning (AL) and domain adaptation (DA). The contributions of this article are three-fold: 1) the domain-discrepancy problem in intelligent logging lithology identification is first investigated in this article, and a novel framework that incorporates AL and DA into lithology identification is proposed to handle the problem; 2) we design a discrepancy-based AL and pseudolabeling (PL) module and an instance importance weighting module to query the most uncertain target information and retain the most confident source information, which solves the challenges of cost limitation and distribution misalignment; and 3) we develop a reliability detecting module to improve the reliability of target pseudolabels, which, together with the discrepancy-based AL and PL module, solves the challenge of data divergence. Extensive experiments on three real-world well-logging datasets demonstrate the effectiveness of the proposed method compared to the baselines. Ji Chang, Yu Kang 0001, Wei Xing Zheng 0001, Yang Cao 0010, Wenjun Lv, Xing-Mou Wang |
IEEE Trans. Cybern. | 4 |
| 2022 | Robust Object Detection via Adversarial Novel Style ExplorationabstractDeep object detection models trained on clean images may not generalize well on degraded images due to the well-known domain shift issue. This hinders their application in real-life scenarios such as video surveillance and autonomous driving. Though domain adaptation methods can adapt the detection model from a labeled source domain to an unlabeled target domain, they struggle in dealing with open and compound degradation types. In this paper, we attempt to address this problem in the context of object detection by proposing a robust object Detector via Adversarial Novel Style Exploration (DANSE). Technically, DANSE first disentangles images into domain-irrelevant content representation and domain-specific style representation under an adversarial learning framework. Then, it explores the style space to discover diverse novel degradation styles that are complementary to those of the target domain images by leveraging a novelty regularizer and a diversity regularizer. The clean source domain images are transferred into these discovered styles by using a content-preserving regularizer to ensure realism. These transferred source domain images are combined with the target domain images and used to train a robust degradation-agnostic object detection model via adversarial domain adaptation. Experiments on both synthetic and real benchmark scenarios confirm the superiority of DANSE over state-of-the-art methods. Wen Wang 0009, Jing Zhang 0037, Wei Zhai, Yang Cao 0010, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2022 | UJ-FLAC: Unsupervised Joint Feature Learning and Clustering for Dynamic Driving Cycles ConstructionabstractDriving cycles construction, which aims to generate various vehicle driving profiles corresponding to typical traffic conditions, plays an important role in the evaluation of vehicle emissions, economy and mileage. Existing methods usually represent the speed-time distributions of driving data in the space spanned by hand-crafted features, and select typical sequences to combine driving cycle curves. However, since the driving data is treated as static, the inherent dynamic characteristics and temporal dependency tend to be ignored, resulting in low accuracy and insufficient robustness. To address this issue, this paper proposes a dynamic driving cycle construction framework, in which feature extraction and sequence clusters are achieved in an unsupervised joint learning manner. Specifically, the driving data are firstly encoded by a Bi-directional Long Short-Term Memory (BiLSTM) branch to capture the temporal correlation property of driving sequences. Then, a temporal clustering branch is presented to achieve soft distribution clustering of feature sequences by introducing a relative-entropy-based regularization term into the coding unit. The two branches are iteratively updated until stable feature learning and clustering results are obtained. Consequently, each branch benefits from the additional improvement over the previous branch during the iteration process. Finally, typical driving sequences are selected according to the intra-class/extra-class distance and class proportion, and then assembled to generate driving cycles profiles. To verify the performance of our proposed method, evaluations are performed on the on-road driving data of light vehicles in Fuzhou, in which the constructed driving cycle from our methods is substituted into COPERT model to estimate and visualize the road emissions, and the experimental results demonstrate that our proposed methods can greatly improve the accuracy and robustness of the constructed driving cycle. Lihong Pei, Yang Cao 0010, Yu Kang 0001, Zhenyi Xu, Zhen-Yi Zhao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Self-Promoted Prototype Refinement for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning is to recognize the new classes given few samples and not forget the old classes. It is a challenging task since representation optimization and prototype reorganization can only be achieved under little supervision. To address this problem, we propose a novel incremental prototype learning scheme. Our scheme consists of a random episode selection strategy that adapts the feature representation to various generated incremental episodes to enhance the corresponding extensibility, and a self-promoted prototype refinement mechanism which strengthens the expression ability of the new classes by explicitly considering the dependencies among different classes. Particularly, a dynamic relation projection module is proposed to calculate the relation matrix in a shared embedding space and leverage it as the factor for bootstrapping the update of prototypes. Extensive experiments on three benchmark datasets demonstrate the above-par incremental performance, outperforming state-of-the-art methods by a margin of 13%, 17% and 11%, respectively. Kai Zhu 0004, Yang Cao 0010, Wei Zhai, Zhengjun Zha |
CVPR | 2 |
| 2021 | Adaptive Channel Attention and Feature Super-Resolution for Remote Sensing Images Spatiotemporal FusionabstractCNN-based Spatiotemporal image fusion (STIF) methods have achieved better performance than traditional researches. However, most CNN-based methods fail to make full use of hierarchical features, and ignore the quality and distribution characteristics of feature maps in fine-grained STIF. In this paper, we propose a network with channel attention and feature super-resolution for STIF (CAFSRNet). First, our method uses the low resolution time-domain changing images as input to extract changes more accurately and simplify computational overhead. Second, channel attention mechanism is introduced into Cross-spatial Resolution Mapping module to make the network pay more attention to informative features. Third, by adding feature super-resolution into the supervision process, we enhance the distribution of feature maps and the quality of mapping results. The qualitative and quantitative experimental results on various datasets demonstrate the superiority of our proposed method over the state-of-the-art methods. Shuai Fang, Siyuan Meng, Yang Cao 0010, Jing Zhang 0037, Weikai Shi |
IGARSS | 3 |
| 2021 | Multi-Label Hyperspectral Classification with Discriminative FeaturesabstractFor hyperspectral classification, mixed pixels are usually the biggest reason to reduce the classification accuracy. To solve the problem, we apply the multi-label classification technique to the hyperspectral image classification. The approach Label-specific Features (LIFT) achieves state-of-the-art performance because the most distinctive features are constructed for each label. Clustering centers play an essential role in label-specific feature conversion. However, LIFT does not consider the relationship between positive and negative instances, and clustering centers on the training dataset is inconsistent with that of the original instances. In this paper, we propose a new algorithm called SMOTE_DFL, which can get better clustering centers through two strategies: 1) The spectral clustering algorithm SIA is introduced to focus on the local structure between positive and negative instances; 2) By oversampling training sample, the clustering center of training sample is close to that of the whole. Extensive experiments are conducted on three datasets. The results validate the superiority of SMOTE_DFL to other algorithms. Shuai Fang, Jing Zhang 0037, Yang Cao 0010, Weikai Shi |
IGARSS | 5 |
| 2021 | Cascade Network for Hyperspectral Image ClassificationabstractConvolutional neural network (CNN) is one of the most powerful tools to deal with computer vision tasks such as hyperspectral image (HSI) classification. While many studies using CNN focus on classification precision, few of them pay attention to the model size and running time. Some studies focus on lightweight neural networks for traditional RGB image processing tasks and achieve fantastic results, but none of them are designed for hyperspectral image processing. In this paper, a novel lightweight neural network designed for hyperspectral image classification is proposed to do fast HSI processing while maintaining high classification precision. The network uses the idea of feature reuse to reduce parameter size and improve convergence. Expansion convolution is adopted to overcome the defects which are brought by parameter reduction. The experiments show that the proposed network has SOTA level classification accuracy while maintaining high processing speed. Shuai Fang, Jing Zhang 0037, Yang Cao 0010, Weikai Shi |
IGARSS | 4 |
| 2021 | One-Shot Affordance DetectionabstractAffordance detection refers to identifying the potential action possibilities of objects in an image, which is an important ability for robot perception and manipulation. To empower robots with this ability in unseen scenarios, we consider the challenging one-shot affordance detection problem in this paper, i.e., given a support image that depicts the action purpose, all objects in a scene with the common affordance should be detected. To this end, we devise a One-Shot Affordance Detection (OS-AD) network that firstly estimates the purpose and then transfers it to help detect the common affordance from all candidate images. Through collaboration learning, OS-AD can capture the common characteristics between objects having the same underlying affordance and learn a good adaptation capability for perceiving unseen affordances. Besides, we build a Purpose-driven Affordance Dataset (PAD) by collecting and labeling 4k images from 31 affordance and 72 object categories. Experimental results demonstrate the superiority of our model over previous representative ones in terms of both objective metrics and visual quality. The benchmark suite is at ProjectPage. Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao |
IJCAI | 4 |
| 2021 | Exploring Sequence Feature Alignment for Domain Adaptive Detection TransformersabstractDetection transformers have recently shown promising object detection results and attracted increasing attention. However, how to develop effective domain adaptation techniques to improve its cross-domain performance remains unexplored and unclear. In this paper, we delve into this topic and empirically find that direct feature distribution alignment on the CNN backbone only brings limited improvements, as it does not guarantee domain-invariant sequence features in the transformer for prediction. To address this issue, we propose a novel Sequence Feature Alignment (SFA) method that is specially designed for the adaptation of detection transformers. Technically, SFA consists of a domain query-based feature alignment (DQFA) module and a token-wise feature alignment (TDA) module. In DQFA, a novel domain query is used to aggregate and align global context from the token sequence of both domains. DQFA reduces the domain discrepancy in global feature representations and object relations when deploying in the transformer encoder and decoder, respectively. Meanwhile, TDA aligns token features in the sequence from both domains, which reduces the domain gaps in local and instance-level feature representations in the transformer encoder and decoder, respectively. Besides, a novel bipartite matching consistency loss is proposed to enhance the feature discriminability for robust object detection. Experiments on three challenging benchmarks show that SFA outperforms state-of-the-art domain adaptive object detection methods. Code has been made available at: https://github.com/encounter1997/SFA. Wen Wang 0009, Yang Cao 0010, Jing Zhang 0037, Fengxiang He, Zhengjun Zha, Yonggang Wen 0001, Dacheng Tao |
ACM Multimedia | 2 |
| 2021 | Deep amended COPERT model for regional vehicle emission prediction
Zhenyi Xu, Yu Kang 0001, Yang Cao 0010 |
Sci. China Inf. Sci. | 3 |
| 2021 | A tri-attention enhanced graph convolutional network for skeleton-based action recognitionabstractAbstract Skeleton‐based action recognition has recently attracted a lot of research interests due to its advantage in computational efficiency. Some recent work building upon Graph Convolutional Networks (GCNs) has shown promising performance in this task by modelling intrinsic spatial correlations between skeleton joints. However, these methods only consider local properties of action sequences in the spatial‐temporal domain, and consequently, are limited in distinguishing complex actions with similar local movements. To address this problem, a novel tri‐attention module (TAM) is proposed to guide GCNs to perceive significant variations across local movements. Specifically, the devised TAM is implemented in three steps: i) A dimension permuting unit is proposed to characterise skeleton action sequences in three different domains: body poses, joint trajectories, and evolving projections. ii) A global statistical modelling unit is introduced to aggregate the first‐order and second‐order properties of global contexts to perceive the significant movement variations of each domain. iii) A fusion unit is presented to integrate the features of these three domains together and leverage as orientation for graph convolution at each layer. Through these three steps, significant‐variation frames, joints, and channels can be enhanced. We conduct extensive experiments on two large‐scale benchmark datasets, NTU RGB‐D and Kinetics‐Skeleton. Experimental results demonstrate that the proposed TAM can be easily plugged into existing GCNs and achieve comparable performance with the state‐of‐the‐art methods. Wei Zhai, Yang Cao 0010 |
IET Comput. Vis. | 3 |
| 2021 | Deep representation-based packetized predictive compensation for networked nonlinear systems
Shaofeng Chen, Yang Cao 0010, Yu Kang 0001, Bingyu Sun |
Neural Comput. Appl. | 2 |
| 2021 | One-Shot Texture Retrieval Using Global Grouping MetricabstractTexture retrieval is widely used in the fields of fashion and e-commerce. This paper presents the problem of one-shot texture retrieval: given an example of a new reference texture, we aim to detect and segment all pixels of the same texture category within an arbitrary image. To address this problem, an OS-TR network is proposed to encode both reference and query images into a texture representation space, and a better comparison is made based on the global grouping information. Because the learned texture representation should be invariant to the spatial layout while preserving the rough semantic concepts, we introduce an adaptive directionality-aware module to finely discriminate the orderless texture details. To make full use of the global context information given only a few examples, we incorporate a grouping-attention mechanism into the relation network, resulting in the per-channel modulation of the local relation features. Extensive experiments on two benchmark datasets (i.e., the DTD and ADE20K dataset) and real scenarios demonstrate that our proposed method can achieve above-par segmentation performance and robust generalization across domains. Kai Zhu 0004, Yang Cao 0010, Wei Zhai, Zhengjun Zha |
IEEE Trans. Multim. | 2 |
| 2021 | Spatiotemporal Graph Convolution Multifusion Network for Urban Vehicle Emission PredictionabstractUrban vehicle emission prediction can help the regulation of vehicle pollution and traffic control. However, it is hard to predict the spatiotemporal variation of vehicle emission because of the spatial interactions and temporal correlations between different road segments as well as the high nonlinearity and complexity of vehicle emission variation. The existing methods solve the problem by splitting the region into standard segments or grids based on conventional deep learning methods, without considering that urban vehicle emission varies by graph-structured traffic road network and depends on many complex external environment factors. To address these issues, a spatiotemporal graph convolution multifusion network (ST-MFGCN) is proposed to leverage the graph structural properties as the inherent connectivity of road network for urban vehicle emission prediction, which can capture the vehicle emission spatiotemporal variation patterns and learn the effects of complex environmental factors. The proposed model consists of three parts: 1) a spatiotemporal graph convolution module to capture spatiotemporal dependencies by merging closeness, period, and trend sequences with temporal convolution as well as graph convolution is introduced to model the spatial dependencies; 2) an external factor component to divide multisource external factors into global and individual external features; and 3) a general fusion component to merge the spatiotemporal patterns and the external features as well as fit the mutation of emission measurement data by multifusion strategy. Finally, the proposed model is evaluated on the practical monitoring data of vehicle emission data in Hefei, and the results demonstrate that our proposed model can predict regional vehicle emissions effectively. Zhenyi Xu, Yu Kang 0001, Yang Cao 0010, Zhijun Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Leveraging Deep Statistics for Underwater Image EnhancementabstractUnderwater imaging often suffers from color cast and contrast degradation due to range-dependent medium absorption and light scattering. Introducing image statistics as prior has been proved to be an effective solution for underwater image enhancement. However, relative to the modal divergence of light propagation and underwater scenery, the existing methods are limited in representing the inherent statistics of underwater images resulting in color artifacts and haze residuals. To address this problem, this article proposes a convolutional neural network (CNN)-based framework to learn hierarchical statistical features related to color cast and contrast degradation and to leverage them for underwater image enhancement. Specifically, a pixel disruption strategy is first proposed to suppress intrinsic colors’ influence and facilitate modeling a unified statistical representation of underwater image. Then, considering the local variation of depth of field, two parallel sub-networks: Color Correction Network (CC-Net) and Contrast Enhancement Network (CE-Net) are presented. The CC-Net and CE-Net can generate pixel-wise color cast and transmission map and achieve spatial-varied color correction and contrast enhancement. Moreover, to address the issue of insufficient training data, an imaging model-based synthesis method that incorporates pixel disruption strategy is presented to generate underwater patches with global degradation consistency. Quantitative and subjective evaluations demonstrate that our proposed method achieves state-of-the-art performance. Yang Wang 0015, Yang Cao 0010, Jing Zhang 0037, Feng Wu 0001, Zhengjun Zha |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Deep Degradation Prior for Low-Quality Image ClassificationabstractState-of-the-art image classification algorithms building upon convolutional neural networks (CNNs) are commonly trained on large annotated datasets of high-quality images. When applied to low-quality images, they will suffer a significant degradation in performance, since the structural and statistical properties of pixels in the neighborhood are obstructed by image degradation. To address this problem, this paper proposes a novel deep degradation prior for low-quality image classification. It is based on statistical observations that, in the deep representation space, image patches with structural similarity have uniform distribution even if they come from different images, and the distributions of corresponding patches in low- and high-quality images have uniform margins under the same degradation condition. Therefore, we propose a feature de-drifting module (FDM) to learn the mapping relationship between deep representations of low- and high- quality images, and leverage it as a deep degradation prior (DDP) for low-quality image classification. Since the statistical properties are independent to image content, deep degradation prior can be learned on a training set of limited images without supervision of semantic labels and served in a form of “plugging-in” module of the existing classification networks to improve their performance on degraded images. Evaluations on the benchmark dataset ImageNet-C demonstrate that our proposed DDP can improve the accuracy of the pre-trained network model by more than 20% under various degradation conditions. Even under the extreme setting that only 10 images from CUB-C dataset are used for the training of DDP, our method improves the accuracy of VGG16 on ImageNet-C from 37% to 55%. Yang Wang 0015, Yang Cao 0010, Zhengjun Zha, Jing Zhang 0037, Zhiwei Xiong |
CVPR | 2 |
| 2020 | Deep Structure-Revealed Network for Texture RecognitionabstractTexture recognition is a challenging visual task since various primitives along with their arrangements can be recognized from a same texture image when perceiving with different contexts. Some recent work building on CNNs exploits orderless aggregating to provide invariance to spatial arrangements. However, these methods ignore the inherent structural property of textures, which is a critical cue for distinguishing and describing texture images in the wild. To address this problem, we propose a novel Deep Structure-Revealed Network (DSR-Net) that leverages spatial dependency among the captured primitives as structural representation for texture recognition. Specifically, a primitive capturing module (PCM) is devised to generate multiple primitives from eight directional spatial contexts, in which deep features are firstly extracted under the constrains of direction map and then encoded based on the similarities of neighborhood. Next, these primitives are associated with a dependence learning module (DLM) to generate structural representation, in which a two-way collaborative relationship strategy is introduced to perceive the spatial dependencies among multiple primitives. At last, the structure-revealed texture representations are integrated with spatial ordered information to achieve real-world texture recognition. Evaluation on the five most challenging texture recognition datasets has demonstrated the superiority of the proposed model against state-of-the-art methods. The structure-revealed performances of DSR-Net are further verified on some extensive experiments, including fine-grained classification and semantic segmentation. Wei Zhai, Yang Cao 0010, Zhengjun Zha, Haiyong Xie 0001, Feng Wu 0001 |
CVPR | 2 |
| 2020 | Pan-Sharpening Based On Parallel Pyramid Convolutional Neural NetworkabstractExisting deep learning-based pan-sharpening methods mainly learn spatial information from a high-resolution (HR) panchromatic (PAN) image for each spectral channel. However, due to the own characteristics of remote sensing image data, the spatial information of PAN image often shows weak correlation with some spectral channel, especially for channels non-overlapped by PAN channel. In this paper, we propose a parallel pyramid network (PPN) for pan-sharpening. First, a three-branch parallel structure is proposed for dealing with PAN image detail, multispectral (MS) images detail and spectral property respectively. Second, pyramid network structure is introduced in two detail branches to solve the problem of weak correlation due to scale difference. Third, the feature level fusion in two detail branches is implemented, which utilizes redundancy between channels to solve detail representation of channels non-overlapped by PAN channel. The qualitative and quantitative experimental results on various data sets demonstrate the superiority of our proposed method over the state-of-the-art methods. Shuai Fang, Jing Zhang 0037, Yang Cao 0010 |
ICIP | 4 |
| 2020 | Maskpan: Mask Prior Guided Network For PansharpeningabstractPansharpening aims to generate the high spatial resolution multispectral (HRMS) images by fusing the spatial and spectral information from the low resolution multispectral (LRMS) images and high resolution panchromatic (PAN) images. Although existing pansharpening methods excel at achieving visual pleasing HRMS, they are limited in providing discriminability for visual tasks. To address this problem, this paper proposes a mask prior guided network (MaskPan) for pansharpening, which incorporates high-level semantic features with low-level detail information to improve the visual discrimination and quality of pansharpened images simultaneously. To make full use of the mask prior, the spatial and spectral features in conjunction with the semantic features are firstly fused in feature domain, and then promoted by an attention mechanism. In addition, the semantic segmentation task is introduced as a new metric to evaluate the visual discrimination of pansharpened images. Experimental results show that the proposed MaskPan can effectively enhance image quality and visual discrimination, thereby improving the pansharpening performance. Xue Rui, Yang Cao 0010, Yu Kang 0001, Rui Ba |
ICIP | 2 |
| 2020 | Deep Inhomogeneous Regularization For Transfer LearningabstractFine-tuning is an effective transfer learning method to achieve ideal performance on target task with limited training data. Some recent works regularize parameters of deep neural networks for better knowledge transfer. However, these methods enforce homogeneous penalties for all parameters, resulting in catastrophic forgetting or negative transfer. To address this problem, we propose a novel Inhomogeneous Regularization (IR) method that imposes a strong regularization on parameters of transferable convolutional filters to tackle catastrophic forgetting and alleviate the regularization on parameters of less transferable filters to tackle negative transfer. Moreover, we use the decaying averaged deviation of parameters from the start point (pre-trained parameters) to accurately measure the transferability of each filter. Evaluation on the three challenging benchmarks datasets has demonstrated the superiority of the proposed model against state-of-the-art methods. Wen Wang 0009, Wei Zhai, Yang Cao 0010 |
ICIP | 3 |
| 2020 | Self-Supervised Tuning for Few-Shot SegmentationabstractFew-shot segmentation aims at assigning a category label to each image pixel with few annotated samples. It is a challenging task since the dense prediction can only be achieved under the guidance of latent features defined by sparse annotations. Existing meta-learning based method tends to fail in generating category-specifically discriminative descriptor when the visual features extracted from support images are marginalized in embedding space. To address this issue, this paper presents an adaptive tuning framework, in which the distribution of latent features across different episodes is dynamically adjusted based on a self-segmentation scheme, augmenting category-specific descriptors for label prediction. Specifically, a novel self-supervised inner-loop is firstly devised as the base learner to extract the underlying semantic features from the support image. Then, gradient maps are calculated by back-propagating self-supervised loss through the obtained features, and leveraged as guidance for augmenting the corresponding elements in the embedding space. Finally, with the ability to continuously learn from different episodes, an optimization-based meta-learner is adopted as outer loop of our proposed framework to gradually refine the segmentation results. Extensive experiments on benchmark PASCAL-5i and COCO-20i datasets demonstrate the superiority of our proposed method over state-of-the-art. Kai Zhu 0004, Wei Zhai, Yang Cao 0010 |
IJCAI | 3 |
| 2020 | Nighttime Dehazing with a Synthetic BenchmarkabstractIncreasing the visibility of nighttime hazy images is challenging because of uneven illumination from active artificial light sources and haze absorbing/scattering. The absence of large-scale benchmark datasets hampers progress in this area. To address this issue, we propose a novel synthetic method called 3R to simulate nighttime hazy images from daytime clear images, which first reconstructs the scene geometry, then simulates the light rays and object reflectance, and finally renders the haze effects. Based on it, we generate realistic nighttime hazy images by sampling real-world light colors from a prior empirical distribution. Experiments on the synthetic benchmark show that the degrading factors jointly reduce the image quality. To address this issue, we propose an optimal-scale maximum reflectance prior to disentangle the color correction from haze removal and address them sequentially. Besides, we also devise a simple but effective learning-based baseline which has an encoder-decoder structure based on the MobileNet-v2 backbone. Experiment results demonstrate their superiority over state-of-the-art methods in terms of both image quality and runtime. Both the dataset and source code will be available at https://github.com/chaimi2013/3R. Jing Zhang 0037, Yang Cao 0010, Zhengjun Zha, Dacheng Tao |
ACM Multimedia | 2 |
| 2020 | Deep Palette-Based Color Decomposition for Image Recoloring with Aesthetic Suggestion
Zhengqing Li, Zhengjun Zha, Yang Cao 0010 |
MMM (1) | 3 |
| 2020 | Emission stations location selection based on conditional measurement GAN data
Zhenyi Xu, Yu Kang 0001, Yang Cao 0010 |
Neurocomputing | 3 |
| 2020 | Deep time-frequency representation and progressive decision fusion for ECG classification
Jing Zhang 0037, Yang Cao 0010, Yuxiang Yang 0001, Xiaobin Xu 0002 |
Knowl. Based Syst. | 3 |
| 2020 | Object affordance detection with relationship-aware network
Yang Cao 0010, Yu Kang 0001 |
Neural Comput. Appl. | 2 |
| 2019 | Deep Multiple-Attribute-Perceived Network for Real-World Texture RecognitionabstractTexture recognition is a challenging visual task as multiple perceptual attributes may be perceived from the same texture image when combined with different spatial context. Some recent works building upon Convolutional Neural Network (CNN) incorporate feature encoding with orderless aggregating to provide invariance to spatial layouts. However, these existing methods ignore visual texture attributes, which are important cues for describing the real-world texture images, resulting in incomplete description and inaccurate recognition. To address this problem, we propose a novel deep Multiple-Attribute-Perceived Network (MAP-Net) by progressively learning visual texture attributes in a mutually reinforced manner. Specifically, a multi-branch network architecture is devised, in which cascaded global contexts are learned by introducing similarity constraint at each branch, and leveraged as guidance of spatial feature encoding at next branch through an attribute transfer scheme. To enhance the modeling capability of spatial transformation, a deformable pooling strategy is introduced to augment the spatial sampling with adaptive offsets to the global context, leading to perceive new visual attributes. An attribute fusion module is then introduced to jointly utilize the perceived visual attributes and the abstracted semantic concepts at each branch. Experimental results on the five most challenging texture recognition datasets have demonstrated the superiority of the proposed model against the state-of-the-arts. Wei Zhai, Yang Cao 0010, Jing Zhang 0037, Zhengjun Zha |
ICCV | 2 |
| 2019 | One-Shot Texture Retrieval with Global Context MetricabstractIn this paper, we tackle one-shot texture retrieval: given an example of a new reference texture, detect and segment all the pixels of the same texture category within an arbitrary image. To address this problem, we present an OS-TR network to encoding both reference patch and query image, leading to achieve texture segmentation towards the reference category. Unlike the existing texture encoding methods that integrate CNN with orderless pooling, we propose a directionality-aware network to capture the texture variations at each direction, resulting in spatially invariant representation. To segment new categories given only few examples, we incorporate a self-gating mechanism into relation network to exploit global context information for adjusting per-channel modulation weights of local relation features. Extensive experiments on benchmark texture datasets and real scenarios demonstrate the above-par segmentation performance and robust generalization across domains of our proposed method. Kai Zhu 0004, Wei Zhai, Zhengjun Zha, Yang Cao 0010 |
IJCAI | 4 |
| 2019 | Progressive Retinex: Mutually Reinforced Illumination-Noise Perception Network for Low-Light Image EnhancementabstractContrast enhancement and noise removal are coupled problems for low-light image enhancement. The existing Retinex based methods do not take the coupling relation into consideration, resulting in under or over-smoothing of the enhanced images. To address this issue, this paper presents a novel progressive Retinex framework, in which illumination and noise of low-light image are perceived in a mutually reinforced manner, leading to noise reduction low-light enhancement results. Specifically, two fully pointwise convolutional neural networks are devised to model the statistical regularities of ambient light and image noise respectively, and to leverage them as constraints to facilitate the mutual learning process. The proposed method not only suppresses the interference caused by the ambiguity between tiny textures and image noises, but also greatly improves the computational efficiency. Moreover, to solve the problem of insufficient training data, we propose an image synthesis strategy based on camera imaging model, which generates color images corrupted by illumination-dependent noises. Experimental results on both synthetic and real low-light images demonstrate the superiority of our proposed approaches against the State-Of-The-Art (SOTA) low-light enhancement methods. Yang Wang 0015, Yang Cao 0010, Zhengjun Zha, Jing Zhang 0037, Zhiwei Xiong, Wei Zhang 0021, Feng Wu 0001 |
ACM Multimedia | 2 |
| 2019 | 3D Layout encoding network for spatial-aware 3D saliency modellingabstractThree‐dimensional (3D) [red, green and blue (RGB) + depth] saliency modelling can help with popular 3D multimedia applications. However, depth images produced from existing 3D devices are often with low quality, e.g. containing noises and holes. In this study, rather than relying on features or predictions directly derived from single depth images, the authors propose to encode deep layout features to facilitate the spatial‐aware saliency prediction. Specifically, they first generate coarse depth‐induced saliency cues which are careless of depth details. Then, to leverage the information of the high‐quality RGB image, they embed both low‐level and high‐level RGB deep features to refine the final prediction. In this way, they take both bottom‐up and top‐down cues together with spatial layout into account and achieve better saliency modelling results. Experiments on five public datasets show the superiority of the proposed method. Yang Cao 0010, Yu Kang 0001, Zhongcheng Yin, Rui Ba |
IET Comput. Vis. | 2 |
| 2019 | PixTextGAN: structure aware text image synthesis for license plate recognitionabstractRapid progress on text image recognition has been achieved with the development of deep‐learning techniques. However, it is still a great challenge to achieve a comprehensive license plate recognition in the real scenes, since there are no publicly available large diverse datasets for the training of deep learning models. This paper aims at synthesising of license plate images with generative adversarial networks (GAN), refraining from collecting a vast amount of labelled data. The authors thus propose a novel PixTextGAN that leverages a controllable architecture that generates specific character structures for different text regions to generate synthetic license plate images with reasonable text details. Specifically, a comprehensive structure‐aware loss function is presented to preserve the key characteristic of each character region and thus to achieve appearance adaption for better recognition. Qualitative and quantitative experiments demonstrate the superiority of authors’ proposed method in text image synthetisation over state‐of‐the‐art GANs. Further experimental results of license plate recognition on ReId and CCPD dataset demonstrate that using the synthesised images by PixTextGAN can greatly improve the recognition accuracy. Shilian Wu, Wei Zhai, Yang Cao 0010 |
IET Image Process. | 3 |
| 2019 | Deep spatiotemporal residual early-late fusion network for city region vehicle emission pollution prediction
Zhenyi Xu, Yang Cao 0010, Yu Kang 0001 |
Neurocomputing | 2 |
| 2019 | Man-machine verification of mouse trajectory based on the random forest modelabstractIdentifying code has been widely used in man-machine verification to maintain network security. The challenge in engaging man-machine verification involves the correct classification of man and machine tracks. In this study, we propose a random forest (RF) model for man-machine verification based on the mouse movement trajectory dataset. We also compare the RF model with the baseline models (logistic regression and support vector machine) based on performance metrics such as precision, recall, false positive rates, false negative rates, F -measure, and weighted accuracy. The performance metrics of the RF model exceed those of the baseline models. Zhenyi Xu, Yu Kang 0001, Yang Cao 0010 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2019 | Deep Convolutional Identifier for Dynamic Modeling and Adaptive Control of Unmanned HelicopterabstractHelicopters are complex high-order and time-varying nonlinear systems, strongly coupling with aerodynamic forces, engine dynamics, and other phenomena. Therefore, it is a great challenge to investigate system identification for dynamic modeling and adaptive control for helicopters. In this paper, we address the system identification problem as dynamic regression and propose to represent the uncertainties and the hidden states in the system dynamic model with a deep convolutional neural network. Particularly, the parameters of the network are directly learned from the real flight data of aerobatic helicopter. Since the deep convolutional model has a good performance for describing the dynamic behavior of the hidden states and uncertainties in the flight process, the proposed identifier manifests strong robustness and high accuracy, even for untrained aerobatic maneuvers. The effectiveness of the proposed method is verified by various experiments with the real-world flight data from the Stanford Autonomous Helicopter Project. Consequently, an adaptive flight control scheme including a deep convolutional identifier and a backstepping-based controller is presented. The stability of the flight control scheme is rigorously proved by the Lyapunov theory. It reveals that the tracking errors for both the position and attitude of unmanned helicopter asymptotic converge to a small neighborhood of the origin. Yu Kang 0001, Shaofeng Chen, Yang Cao 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | A Generative Adversarial Network Based Framework for Unsupervised Visual Surface InspectionabstractVisual surface inspection is a challenging task due to the highly inconsistent appearance of the target surfaces and the abnormal regions. Most of the state-of-the-art methods are highly dependent on the labelled training samples, which are difficult to collect in practical industrial applications. To address this problem, we propose a generative adversarial network based framework for unsupervised surface inspection. The generative adversarial network is trained to generate the fake images analogous to the normal surface images. It implies that a well-trained GAN indeed learns a good representation of the normal surface images in a latent feature space. And consequently, the discriminator of GAN can serve as a naturally one-class classifier. We use the first three conventional layer of the discriminator as the feature extractor, whose response is sensitive to the abnormal regions. Particularly, a multi-scale fusion strategy is adopted to fuse the responses of the three convolution layers and thus improve the segmentation performance of abnormal detection. Various experimental results demonstrate the effectiveness of our proposed method. Wei Zhai, Yang Cao 0010, Zengfu Wang |
ICASSP | 3 |
| 2018 | The Deep Input-Koopman Operator for Nonlinear Systems
Rongrong Zhu, Yang Cao 0010, Yu Kang 0001 |
ICONIP (7) | 2 |
| 2018 | Fully Point-wise Convolutional Neural Network for Modeling Statistical Regularities in Natural ImagesabstractModeling statistical regularity plays an essential role in ill-posed image processing problems. Recently, deep learning based methods have been presented to implicitly learn statistical representation of pixel distributions in natural images and leverage it as a constraint to facilitate subsequent tasks, such as color constancy and image dehazing. However, the existing CNN architecture is prone to variability and diversity of pixel intensity within and between local regions, which may result in inaccurate statistical representation. To address this problem, this paper presents a novel fully point-wise CNN architecture for modeling statistical regularities in natural images. Specifically, we propose to randomly shuffle the pixels in the origin images and leverage the shuffled image as input to make CNN more concerned with the statistical properties. Moreover, since the pixels in the shuffled image are independent identically distributed, we can replace all the large convolution kernels in CNN with point-wise (1*1) convolution kernels while maintaining the representation ability. Experimental results on two applications: color constancy and image dehazing, demonstrate the superiority of our proposed network over the existing architectures, i.e., using 1/10~1/100 network parameters and computational cost while achieving comparable performance. Jing Zhang 0037, Yang Cao 0010, Yang Wang 0015, Chenglin Wen, Chang Wen Chen |
ACM Multimedia | 2 |
| 2018 | LA-Net: Layout-Aware Dense Network for Monocular Depth EstimationabstractDepth estimation from monocular images is an ill-posed and inherently ambiguous problem. Recently, deep learning technique has been applied for monocular depth estimation seeking data-driven solutions. However, most existing methods focus on pursuing the minimization of average depth regression error at pixel level and neglect to encode the global layout of scene, resulting in layout-inconsistent depth map. This paper proposes a novel Layout-Aware Convolutional Neural Network (LA-Net) for accurate monocular depth estimation by simultaneously perceiving scene layout and local depth details. Specifically, a Spatial Layout Network (SL-Net) is proposed to learn a layout map representing the depth ordering between local patches. A Layout-Aware Depth Estimation Network (LDE-Net) is proposed to estimate pixel-level depth details using multi-scale layout maps as structural guidance, leading to layout-consistent depth map. A dense network module is used as the base network to learn effective visual details resorting to dense feed-forward connections. Moreover, we formulate an order-sensitive softmax loss to well constrain the ill-posed depth inferring problem. Extensive experiments on both indoor scene (NYUD-v2) and outdoor scene (Make3D) datasets have demonstrated that the proposed LA-Net outperforms the state-of-the-art methods and leads to faithful 3D projections. Kecheng Zheng, Zhengjun Zha, Yang Cao 0010, Xuejin Chen, Feng Wu 0001 |
ACM Multimedia | 3 |
| 2018 | Crowd Distribution Estimation with Multi-scale Recursive Convolutional Neural Network
Yu Kang 0001, Yang Cao 0010 |
MMM (1) | 4 |
| 2018 | Co-occurrent Structural Edge Detection for Color-Guided Depth Map Super-Resolution
Wei Zhai, Yang Cao 0010, Zhengjun Zha |
MMM (1) | 3 |
| 2017 | Fast Haze Removal for Nighttime Image Using Maximum Reflectance PriorabstractIn this paper, we address a haze removal problem from a single nighttime image, even in the presence of varicolored and non-uniform illumination. The core idea lies in a novel maximum reflectance prior. We first introduce the nighttime hazy imaging model, which includes a local ambient illumination item in both direct attenuation term and scattering term. Then, we propose a simple but effective image prior, maximum reflectance prior, to estimate the varying ambient illumination. The maximum reflectance prior is based on a key observation: for most daytime haze-free image patches, each color channel has very high intensity at some pixels. For the nighttime haze image, the local maximum intensities at each color channel are mainly contributed by the ambient illumination. Therefore, we can directly estimate the ambient illumination and transmission map, and consequently restore a high quality haze-free image. Experimental results on various nighttime hazy images demonstrate the effectiveness of the proposed approach. In particular, our approach has the advantage of computational efficiency, which is 10-100 times faster than state-of-the-art methods. Jing Zhang 0037, Yang Cao 0010, Shuai Fang, Yu Kang 0001, Chang Wen Chen |
CVPR | 2 |
| 2017 | A deep CNN method for underwater image enhancementabstractUnderwater images often suffer from color distortion and visibility degradation due to the light absorption and scattering. Existing methods utilize various assumptions/constrains to achieve reasonable solutions for underwater image enhancement. However, these methods share the common limitation that the adopted assumptions may not work for some particular scenes. To address this problem, this paper proposes an end to end framework for underwater image enhancement, where a CNN-based network called UIE-Net is presented. The UIE-net is trained with two tasks, color correction and haze removal. This unified training approach enables learning a strong feature representation for both tasks simultaneously. For better extracting the inherent features in local patches, a pixels disrupting strategy is exploited in the proposed learning framework, which significantly improves the convergent speed and accuracy. To handle the training of UIE-net, we synthesize 200000 training images based on the physical underwater imaging model. Experiments on benchmark underwater images for cross-scenes show that UIE-net achieves superior performance over existing methods. Yang Wang 0015, Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 3 |
| 2017 | Image guided depth enhancement via deep fusion and local linear regularizaronabstractDepth maps captured by RGB-D cameras are often noisy and incomplete at edge regions. Most existing methods assume that there is a co-occurrence of edges in depth map and its corresponding color image, and improve the quality of depth map guided by the color image. However, when the color image is noisy or richly detailed, the high frequency artifacts will be introduced into depth map. In this paper, we propose a deep residual network based on deep fusion and local linear regularization for guided depth enhancement. The presented scheme can effectively extract the correlation between depth map and color image in the deep feature space. To reduce the difficulty of training, a specific layer of network which introduces a local linear regularization constraint on the output depth is designed. Experiments on various applications, including depth denoising, super-resolution and inpainting, demonstrate the effectiveness and reliability of our proposed approach. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 3 |
| 2017 | Deep CNN Identifier for Dynamic Modelling of Unmanned Helicopter
Shaofeng Chen, Yang Cao 0010, Yu Kang 0001, Rongrong Zhu, Pengfei Li 0006 |
ICONIP (6) | 2 |
| 2017 | Packet-Dropouts Compensation for Networked Control System via Deep ReLU Neural Network
Yang Cao 0010, Yu Kang 0001, Pengfei Li 0006 |
ICONIP (6) | 2 |
| 2017 | A networked remote sensing system for on-road vehicle emission monitoring
Yu Kang 0001, Yang Cao 0010, Yun-Bo Zhao |
Sci. China Inf. Sci. | 4 |
| 2017 | Simultaneously retargeting and super-resolution for stereoscopic video
Yang Cao 0010, Zengfu Wang |
Multim. Tools Appl. | 2 |
| 2016 | Fast depth estimation from single image using structured forestabstractDepth estimation from single image is an important component of many vision systems, including robot navigation, motion capture and video surveillance. In this paper, we propose to apply a structure forest framework to infer depth information from single RGB image. The core idea of our approach is to exploit the structure properties exhibit in local patches of depth map to learn the depth level for each pixel. We formulate the problem of depth estimation in a structured learning framework based on random decision forests. Each trained forest infers a patch of structured labels that are accumulated across the image to obtain the final depth map. Moreover, we systematically investigate a variety of depth-relevant features and the regression forest framework automatically determines the best feature combination and uses as input of structure forest. Our approach achieves quasi real-time performance that is orders of magnitude faster than state-of-the-art approaches, while also achieving state-of-the-art depth estimation results on the Make3D dataset. Shuai Fang, Ren Jin, Yang Cao 0010 |
ICIP | 3 |
| 2016 | An iterative method for optical flow estimation with motion blurabstractThis paper presents a new method for estimating the optical flow of image sequences while considering the blur effect. In fact, the blur in input images can degrade the quality of optical flow, because it leads to ambiguities of pixel-match. Our method begins with an initial optical flow. Then two steps are performed iteratively until convergence, 1) the blur kernel is estimated using the information from optical flow; 2) the optical flow is estimated considering the blur kernel. Various experimental results verify the effectiveness of our method. Xiangxi Shi, Yang Cao 0010 |
VCIP | 3 |
| 2016 | Automatic chessboard corner detection methodabstractChessboard corner detection is a necessary procedure of the popular chessboard pattern‐based camera calibration technique, in which the inner corners on a two‐dimensional chessboard are employed as calibration markers. In this study, an automatic chessboard corner detection algorithm is presented for camera calibration. In authors’ method, an initial corner set is first obtained with an improved Hessian corner detector. Then, a novel strategy that utilises both intensity and geometry characteristics of the chessboard pattern is presented to eliminate fake corners from the initial corner set. After that, a simple yet effective approach is adopted to sort the detected corners into a meaningful order. Finally, the sub‐pixel location of each corner is calculated. The proposed algorithm only requires a user input of the chessboard size, while all the other parameters can be adaptively calculated with a statistical approach. The experimental results demonstrate that the proposed method has advantages over the popular OpenCV chessboard corner detection method in terms of detection accuracy and computational efficiency. Furthermore, the effectiveness of the proposed method used for camera calibration is also verified in authors’ experiments. Yu Liu 0023, Shuping Liu, Yang Cao 0010, Zengfu Wang |
IET Image Process. | 3 |
| 2016 | Automatic tag saliency ranking for stereo images
Yang Cao 0010, Jing Zhang 0037, Zengfu Wang |
Neurocomputing | 1 |
| 2016 | Salient object detection and classification for stereoscopic images
Yang Cao 0010, Jing Zhang 0037, Zengfu Wang |
Multim. Tools Appl. | 2 |
| 2016 | A Unified Scheme for Super-Resolution and Depth Estimation From Asymmetric Stereoscopic VideoabstractReconstructing a full-resolution stereoscopic video from an asymmetric stereoscopic video is a challenging task. The existing approaches require depth information, which imposes an additional challenge in data acquisition. In this paper, we propose a novel scheme that is capable of obtaining super-resolution and depth estimation simultaneously from an asymmetric stereoscopic video. The proposed scheme models the video super-resolution and stereo matching with a unified energy function. Then, we apply an alternating optimization method to minimize this energy function, which can be implemented with a two-step algorithm. In the first step we calculate the initial depth map by using a region-based cooperative optimization technique while considering the temporal consistency in video. In the second step we resolve the super-resolution problem under the guidance of the depth information. It is effective because each step benefits from the additional improvement over the previous step. We iteratively update the two steps until stable depth and super-resolution results are obtained. We have conducted a series of experiments on public stereoscopic video sequences to evaluate the performance of the proposed method. Both objective indexes and subjective visual comparisons verify that the proposed scheme can achieve satisfactory super-resolution results and high-quality depth map simultaneously. In particular, the subjective evaluation experiments on a 3-D monitor show that this scheme outperforms others and achieves the best visual sharpness. Jing Zhang 0037, Yang Cao 0010, Zhengjun Zha, Zhigang Zheng, Chang Wen Chen, Zengfu Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | A new image filtering method: Nonlocal image guided averagingabstractImage guided filtering has been widely used in many image processing applications. However, it is a local filtering method and has limited propagation ability. In this paper, we propose a new image filtering method: nonlocal image guided averaging (NLGA). Derived from a nonlocal linear model, the proposed method can utilize the nonlocal similarity of the guidance image, so that it can propagate nonlocal information reliably. Consequently, NLGA can obtain a sharper filtering results in the edge regions and more smooth results in the smooth regions. It shows superiority over image guided filtering in different applications, such as image dehazing, depth map super-resolution and image denoising. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICASSP | 2 |
| 2014 | A practical algorithm for automatic chessboard corner detectionabstractChessboard corner detection is a fundamental work of the popular chessboard pattern-based camera calibration technique. In this paper, a fast and robust algorithm for chessboard corner detection is presented. In our method, an initial corner set is obtained with an improved Hessian corner detector. And then, a novel strategy which takes both textural and geometrical characteristics of a chessboard into consideration is employed to eliminate fake corners in the initial corner set. The proposed algorithm only requires a user-input of the total number of chessboard inner corners, while all the other parameters can be adaptively calculated with a statistical approach. Experimental results on two public data sets demonstrate that the proposed method can outperform the most commonly used OpenCV method in terms of both detection rate and computational efficiency. Yu Liu 0023, Shuping Liu, Yang Cao 0010, Zengfu Wang |
ICIP | 3 |
| 2014 | Nighttime haze removal based on a new imaging modelabstractNighttime haze removal is important for different applications such as nighttime video surveillance in haze environment. Different from the imaging conditions in the daytime, nighttime haze images may suffer from non-uniform illumination from artificial light sources. In this paper, we proposed a novel efficient dehazing method with illumination estimation for nighttime haze condition. First, we estimate the light intensity and enhance it to obtain an illumination balanced result. Then, we process a color correction step after estimating the color characteristics of the incident light. Finally, we remove the haze by using the dark channel prior along with estimating the pointwise environmental light. Experimental results show that the proposed method can achieve both illumination balanced and haze free results. Moreover, it also has good color rendition ability. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 2 |
| 2014 | Underwater stereo image enhancement using a new physical modelabstractStereo image applications are becoming more and more prevalent. However, there has been little research on stereo image enhancement. In this paper, we address the challenging problem of underwater stereo image enhancement. A new underwater imaging model is proposed and it can better describe the degradation of underwater images including color distortion and contrast attenuation. In addition, a novel observation that the intensity of the water part within the image is mainly contributed by the scattering light is also proposed. Coupling the proposed model and prior together, the parameters of scattering light can be estimated. Then an iterative approach to process stereo matching and stereo image enhancement alternatively is presented, which can significantly improve the quality of the images and depth maps. The experimental results demonstrate that the proposed method can significantly enhance the image visibility and achieve better depth perception. Jing Zhang 0037, Shuai Fang, Yang Cao 0010 |
ICIP | 4 |
| 2014 | A novel segmentation based video-denoising method with noise level estimation
Yang Cao 0010, Zhengjun Zha, Jing Zhang 0037, Chang Wen Chen |
Inf. Sci. | 1 |
| 2014 | A new closed loop method of super-resolution for multi-view images
Jing Zhang 0037, Yang Cao 0010, Zhigang Zheng, Chang Wen Chen, Zengfu Wang |
Mach. Vis. Appl. | 2 |
| 2013 | A simultaneous method for 3D video super-resolution and high-quality depth estimationabstractMixed-resolution approach serves as a feasible solution to 3D video data reduction in limited bandwidth network environments, i.e., mobile networks. In this paper, we propose a simultaneous method for video super-resolution and high-quality depth estimation of mixed-resolution 3D video. Our method tackles the problem in a joint manner: i). Depth Estimation. We calculate the initial depths by stereo matching, and then warp them according to the optical flow field. ii). Video Super-resolution. Under the guidance of the warped depth information, we resolve the super-resolution problem by a fusion method, which involves a mapping step and a nonlocal reconstruction step. We run the above two steps iteratively until obtaining the stable depth and super-resolution result. The experimental results on public 3D video sequences verify the effectiveness of our proposed method. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 2 |
| 2013 | 2D/3D Model-Based Facial Video Coding/Decoding at Ultra-Low Bit-Rate
Jun Yu 0001, Zengfu Wang, Yang Cao 0010 |
MMM (2) | 3 |
| 2013 | A New Closed Loop Method of Super-Resolution for Multi-view Images
Jing Zhang 0037, Yang Cao 0010, Zhigang Zheng, Zengfu Wang |
MMM (2) | 2 |
| 2013 | A Novel Segmentation-Based Video Denoising Method with Noise Level Estimation
Jing Zhang 0037, Shuai Fang, Yang Cao 0010 |
MMM (1) | 5 |
| 2013 | Digital Multi-Focusing From a Single Photograph Taken With an Uncalibrated Conventional CameraabstractThe demand to restore all-in-focus images from defocused images and produce photographs focused at different depths is emerging in more and more cases, such as low-end hand-held cameras and surveillance cameras. In this paper, we manage to solve this challenging multi-focusing problem with a single image taken with an uncalibrated conventional camera. Different from all existing multi-focusing approaches, our method does not need to include a deconvolution process, which is quite time-consuming and will cause ringing artifacts in the focused region and low depth-of-field. This paper proposes a novel systematic approach to realize multi-focusing from a single photograph. First of all, with the optical explanation for the local smooth assumption, we present a new point-to-point defocus model. Next, the blur map of the input image, which reflects the amount of defocus blur at each pixel in the image, is estimated by two steps. 1) With the sharp edge prior, a rough blur map is obtained by estimating the blur amount at the edge regions. 2) The guided image filter is applied to propagate the blur value from the edge regions to the whole image by which a refined blur map is obtained. Thus far, we can restore the all-in-focus photograph from a defocused input. To further produce photographs focused at different depths, the depth map from the blur map must be derived. To eliminate the ambiguity over the focal plane, user interaction is introduced and a binary graph cut algorithm is used. So we introduce user interaction and use a binary graph cut algorithm to eliminate the ambiguity over the focal plane. Coupled with the camera parameters, this approach produces images focused at different depths. The performance of this new multi-focusing algorithm is evaluated both objectively and subjectively by various test images. Both results demonstrate that this algorithm produces high quality depth maps and multi-focusing results, outperforming the previous approaches. Yang Cao 0010, Shuai Fang, Zengfu Wang |
IEEE Trans. Image Process. | 1 |
| 2011 | Single Image Multi-focusing Based on Local Blur EstimationabstractIn this paper, we address a challenging problem of multi-focusing image from a single photograph taken with an uncalibrated conventional camera. In order to achieve this, we firstly derive an optical degradation model which enables us to adopt a point operation scheme to realize image multi-focusing. This scheme can effectively reduce halo artifacts in the refocused image and greatly improve the computational efficiency. Then, a two-step approach is applied to estimate the blur map of the input image. i). A sparse blur map is obtained by estimating the amount of defocus blur at edge locations. ii). The guided image filtering method is applied to propagate the value from edge locations into the unknown regions. In order to obtain the depth map of the whole scene to realize the multi-focusing, we adopt a simple geometry prior of photograph to eliminate the ambiguity over the focal plane. Based on the obtained depth map, we can directly produce different styles of images by multi-focusing with the adjustment to the camera parameters. Experimental results on a variety of images show that our method can acquire visual pleasing multi-focusing results. Moreover, our method can also extract the depth map of the scene with fairly good extent of accuracy. Yang Cao 0010, Shuai Fang |
ICIG | 1 |
| 2010 | A Close-Form Iterative Algorithm for Depth Inferring from a Single Image
Yang Cao 0010, Zengfu Wang |
ECCV (5) | 1 |
| 2010 | Improved single image dehazing using segmentationabstractIn the hazy weather, the image of outdoor scene is degraded by suspended particles. Scattering and absorption hinder scene radiance and bring in environment light into camera. In this work, a novel algorithm is introduced to restore the clear day image by the segmented hazy image. First, the existing visibility restoration model is analyzed and a conclusion is drawn that the model will violate the contrast enhancement constraint in some specific situations. Next, the graph-based image segmentation method is applied to segment the hazed image by choosing the optimal parameter. Then, the transmission maps prior are obtained according to the blackbody theory. After that, a bilateral filter is designed to amend the transmission map, which can make up the deficiency of restoration model and ensure the transmission map smooth under the contrast enhancement constraint. Last, the experimental results show that the method achieves rather good dehazing results. Shuai Fang, Jiqing Zhan, Yang Cao 0010, Ruizhong Rao |
ICIP | 3 |