Xin Yang 0011

dblp:44/1152-11 · DBLP profile ↗
← Back
119ranked-venue papers
15as first author
85since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 82 · 13 first-author · 54 since 2021Artificial intelligence and machine learning · 59 · 4 first-author · 42 since 2021Systems, architecture and hardware · 6 · 5 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 View-on-Graph: Zero-Shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
abstract
3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision–language models (VLMs) by converting 3D spatial information (SI) into forms amenable to VLM processing, typically as composite inputs such as specified-view renderings or video sequences with overlaid object markers. However, this VLM ⊕ SI paradigm yields entangled visual representations that compel the VLM to process entire cluttered cues, making it hard to exploit spatial–semantic relationships effectively. In this work, we propose a new VLM ⊗ SI paradigm that externalizes the 3D SI into a form enabling the VLM to incrementally retrieve only what it needs during reasoning. We instantiate this paradigm with a novel View-on-Graph (VoG) method, which organizes the scene into a multi-modal, multi-layer scene graph and allows the VLM to operate as an active agent that selectively accesses necessary cues as it traverses the scene. This design offers two intrinsic advantages: (i) by structuring 3D context into a spatially and semantically coherent scene graph rather than confounding the VLM with densely entangled visual inputs, it lowers the VLM's reasoning difficulty; and (ii) by actively exploring and reasoning over the scene graph, it naturally produces transparent, step-by-step traces for interpretable 3DVG. Extensive experiments show that VoG achieves state-of-the-art zero-shot performance, establishing structured scene exploration as a promising strategy for advancing zero-shot 3DVG.
Haiyang Mei, Dongyang Zhan, Jiayue Zhao, Bo Dong 0004, Xin Yang 0011
AAAI7
2026 AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual Tracking
Jiqing Zhang, Yang Wang 0106, Yuanchen Wang, Xin Yang 0011
AAAI7
2026 Dynamic Weight Adaptation in Spiking Neural Networks Inspired by Biological Homeostasis
abstract
Homeostatic mechanisms play a crucial role in maintaining optimal functionality within the neural circuits of the brain. By regulating physiological and biochemical processes, these mechanisms ensure the stability of an organism’s internal environment, enabling it to better adapt to external changes. Among these mechanisms, the Bienenstock, Cooper, and Munro (BCM) theory has been extensively studied as a key principle for maintaining the balance of synaptic strengths in biological systems. Despite the extensive development of spiking neural networks (SNNs) as a model for bionic neural networks, no prior work in the machine learning community has integrated biologically plausible BCM formulations into SNNs to provide homeostasis. In this study, we propose a Dynamic Weight Adaptation Mechanism (DWAM) for SNNs, inspired by the BCM theory. DWAM can be integrated into the host SNN, dynamically adjusting network weights in real time to regulate neuronal activity, providing homeostasis to the host SNN without any fine-tuning. We validated our method through dynamic obstacle avoidance and continuous control tasks under both normal and specifically designed degraded conditions. Experimental results demonstrate that DWAM not only enhances the performance of SNNs without existing homeostatic mechanisms under various degraded conditions but also further improves the performance of SNNs that already incorporate homeostatic mechanisms.
Yunduo Zhou, Bo Dong 0004, Yuanchen Wang, Xuefeng Yin, Yang Wang 0106, Xin Yang 0011
AAAI7
2026 A Review of Learning Based Visual Relocalization Methods
abstract
In recent years, visual relocalization has emerged as a pivotal task in the domains of 3D computer vision and learning based methodologies, witnessing substantial advances due to the evolution of learning based methodologies. This article reviews the current landscape and recent developments in visual relocalization research. It methodically discusses visual relocalization tasks, delineates fundamental solution methodologies, categorizes existing studies, and outlines research objectives within this field. By systematically organizing and elucidating the state-of-the-art of visual relocalization through the lens of learning based methodologies, this paper aims to provide a comprehensive analysis to aid researchers to swiftly grasp the essence of the research problem. It offers a lucid overview of the specific advances in various research directions, thereby facilitating effective applications and further investigation. Additionally, this article anticipates future research trajectories to address visual relocalization challenges.
Xin Yang 0011
Comput. Vis. Media4
2026 You Only Look Intensity Once: Event-Driven Long-Term High-Speed Object Detection
Wen Dong 0008, Haiyang Mei, Yinglian Ji, Ziqi Wei 0001, Shengfeng He, Xin Yang 0011
Int. J. Comput. Vis.7
2026 Efficient Vision Transformer with Token Sparsification for Event-Based Object Tracking
Jiqing Zhang, Xin Yang 0011, Haoming Tang, Yuanchen Wang, Huibing Wang, Xianping Fu
Int. J. Comput. Vis.2
2026 Improving large models with small models: Lower costs and better performance
Dong Chen 0017, Shuo Zhang 0014, Yueting Zhuang, Siliang Tang, Qidong Liu 0001, Xin Yang 0011, Mingliang Xu 0001
Neural Networks8
2026 ScaleGraph: A scalable self-supervised framework for cross-domain zero-shot graph learning
Youjiang Fang, Liang Zhang 0031, Ziqi Wei 0001, Zhichao Wu 0001, Chuanbin Liu 0003, Xin Yang 0011
Pattern Recognit.7
2026 A Lightweight Polarization-Guided Plug-In for Underwater Image Enhancement
abstract
Underwater images play a vital role in marine exploration, but are often severely degraded due to complex imaging conditions, including color distortion, haze effects, and non-uniform illumination. Existing deep learning-based enhancement methods predominantly rely on conventional RGB sensors, which struggle to distinguish between scattered and reflected light, thereby limiting enhancement performance. Polarization imaging, with its capability to capture directional light information, offers promising potential for underwater image enhancement. In this paper, we propose a lightweight yet effective polarization feature extractor that captures global spatial cues from polarization images. Additionally, we design a polarization-guided feature integration module that adaptively enhances the representational capacity of RGB features. Notably, the proposed module is plug-in and can be seamlessly integrated into existing RGB-based enhancement networks. Extensive experiments across multiple datasets demonstrate that incorporating polarization information significantly improves enhancement performance, highlighting its effectiveness as a valuable cue for underwater image enhancement. The code and pretrained models are at https://github.com/jgy0/UPGD.
Guangyao Ju, Jiqing Zhang, Jingqi Zang, Zetian Mi, Xin Yang 0011, Huibing Wang, Jiarui Fan, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.7
2026 Lightweight and Personalized Single-Eye Emotion Recognition via CNN-SNN Spatiotemporal Learning and Memory-Inferred Event Features
abstract
Emotion recognition is essential for improving user experience and interaction quality in human-centered applications. While recent studies have leveraged both event and traditional cameras to enhance eye-based emotion recognition, their practical deployment is hindered by the scarcity of event cameras and the complexity of dual-modality frameworks. Personalization, which is critical for handling individual differences in emotional expression, is also affected by these factors, resulting in reduced performance and adaptation efficiency. To address these challenges, we propose a lightweight and personalized single-eye emotion recognition network, called LPSEER. LPSEER introduces a novel hybrid neural architecture that integrates a convolutional neural network (CNN) and a spiking neural network (SNN) to capture spatiotemporal features from video frames and events, respectively. Additionally, we design a memorybased event feature inference (MEFI) module that recalls event features from video frames, eliminating the reliance on event cameras during inference and personalization while retaining the discriminative advantages of event-based representations. Experimental results demonstrate that LPSEER achieves state-of- the-art recognition accuracy while maintaining the smallest model size and lowest computational cost. Further experiments confirm the strong generalization capabilities and the ability to achieve faster, more accurate personalization. These advantages collectively enable lightweight, accurate, and efficient emotion recognition for real-world human-centered applications.
Qianhui Liu, Jiqing Zhang, Yang Wang 0106, Malu Zhang, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 ICAD: Rethinking the Role of Inference and Cues for the Anomaly Detection of Time Series in IIoT
abstract
Time series anomaly detection in real-world Industrial Internet of Things (IIoT) systems is pivotal for identifying unsafe conditions and implementing timely preventive measures. While diffusion models are popular for capturing complex patterns, they often struggle to balance diversity and fidelity across scenarios due to limited exploration of context-window logical inference relationships and trend-pattern cues. To address these challenges, we propose ICAD, a novel method that rethinks the role of inference and cues in IIoT time series anomaly detection. ICAD defines trend patterns and uncertainties in textual form and utilizes a fine-tuned large language model to encode these descriptions as conditions for the diffusion model, thereby enhancing its generalization across diverse data distributions. Additionally, a “reasoning network for contextual window” mechanism is designed to capture temporal dependencies between adjacent windows, complemented by multi-scale and spatial feature adaptive fusion modules to further enhance the predictive performance. Empirical evaluations across four benchmark datasets and a large-scale ethylene oxide production process demonstrate that ICAD consistently outperforms state-of-the-art baselines, confirming its effectiveness and practicality in overcoming current anomaly detection model limitations.
Zhichao Wu 0001, Zitao Yin, Shunqi Zhang, Xirong Xu, Xiaopeng Wei, Xin Yang 0011
IEEE Trans. Knowl. Data Eng.7
2026 Temporal Prompt Learning With Depth Memory for Video Mirror Detection
abstract
Mirror detection in dynamic scenes plays a crucial role in ensuring safety for various applications, such as drone tracking and robot navigation. However, current mirror detection models often fail in areas with mirrors that have a similar visual and color appearance to their surrounding objects. They also struggle to generalize well in complex cases, primarily due to limited annotated datasets. In this work, we propose a novel temporal prompt learning network with depth memory (TPD-Net) to address these critical challenges. Our approach includes several key components. First, we introduce a Temporal Prompt Generator (TPG) to learn temporal prompt features. Then, we devise Multi-layer Depth-aware Adaptor (MDA) modules to progressively adapt prompt features from the TPG, thereby learning mirror-related features by embedding temporal depth information as guidance. Moreover, we further refine these mirror-related features by constructing a depth memory and a Depth Memory Read module to read the temporal depths stored in the memory, boosting video mirror detection. Experimental results on a benchmark dataset show that our TPD-Net significantly outperforms 22 state-of-the-art methods in video mirror detection tasks. Our code, models, and results are publicly available athttps://github.com/ge-xing/TPDNet.
Zhaohu Xing, Tian Ye 0001, Xin Yang 0011, Sixiang Chen, Huazhu Fu, Yan Nei Law, Lei Zhu 0003
IEEE Trans. Multim.3
2025 Localization Hints Exploration for Object Matting
abstract
Most existing researches achieve image matting via predefined trimap or saliency-guided intermediate segmentation variants, and their essence is providing localization hints for the matting targets. In this paper, we propose our Localization Hints Exploration object matting model (LHEMatt) to estimate alpha mattes based on simple and flexible localization hints. We design our location constraints enhancement module (LCE) to exploit more efficient and effective hints for alpha matte prediction. A well-designed foreground details extension module (FDE) is also presented to integrate target locations with rich boundary details. Besides, to accommodate more computer vision tasks and practical applications, we extend the matting scope to the multi-object foreground and explore separated localization hints to disassemble several targets for object-level matting. Based on this motivation, we construct a new dataset consisting of 757 different delicate alpha mattes (Mobjects-757), and most of them contain two or more objects. As the first object matting inspiration, we perform extensive experiments to demonstrate the proposed localization hints exploration model and evaluate different approaches on the newly-created Mobjects-757 matting dataset.
Yu Qiao 0001, Tianyu Meng, Huilin Ge, Jiayue Zhao, Qianchen Xia, Xin Yang 0011
ICME7
2025 Can I Trust You? Advancing GUI Task Automation with Action Trust Score
Haiyang Mei, Difei Gao, Xiaopeng Wei, Xin Yang 0011, Zheng Shou 0001
ACM Multimedia4
2025 Eye-based Emotion Recognition via Event-Driven Sparse Transformers
abstract
Event-driven eye-based emotion recognition has attracted increasing attention due to the high temporal resolution and dynamic range inherent to event cameras. The intrinsic spatial sparsity of event data, combined with the eye-based emotion recognition task's reliance on localized features such as eyebrows and eyelids, makes it intuitive and efficient to discard less informative regions. However, integrating such sparsification into CNNs remains challenging due to their reliance on dense grid-based operations. In this paper, we propose an efficient vision transformer framework for eye-based emotion recognition with event cameras. Specifically, we present window selection and token selection schemes tailored for event data and eye-based emotion recognition, which can diminish computing demands while enhancing performance. Firstly, we estimate the importance of all local windows and discard those with limited information, reducing computational cost while emphasizing attention on the periocular region. Secondly, we further introduce an adaptive token pruning mechanism that jointly evaluates the input event data and tokens to predict a binary decision mask, identifying and discarding uninformative tokens. Extensive experiments validate that the proposed approach outperforms existing state-of-the-art methods in accuracy by a significant margin.
Zixuan Wan, Jiqing Zhang, Yafei Wang 0004, Zetian Mi, Xin Yang 0011, Xianping Fu, Huibing Wang
ACM Multimedia7
2025 Fully Autonomous Neuromorphic Navigation and Dynamic Obstacle Avoidance
abstract
Unmanned aerial vehicles could accurately accomplish complex navigation and obstacle avoidance tasks under external control. However, enabling unmanned aerial vehicles (UAVs) to rely solely on onboard computation and sensing for real-time navigation and dynamic obstacle avoidance remains a significant challenge due to stringent latency and energy constraints. Inspired by the efficiency of biological systems, we propose a fully neuromorphic framework achieving end-to-end obstacle avoidance during navigation with an overall latency of just 2.3 milliseconds. Specifically, our bio-inspired approach enables accurate moving object detection and avoidance without requiring target recognition or trajectory computation. Additionally, we introduce the first monocular event-based pose correction dataset with over 50,000 paired and labeled event streams. We validate our system on an autonomous quadrotor using only onboard resources, demonstrating reliable navigation and avoidance of diverse obstacles moving at speeds up to 10 m/s.
Xiaochen Shang, Pengwei Luo, Jiayue Zhao, Huilin Ge, Bo Dong 0004, Xin Yang 0011
NeurIPS7
2025 FutuTP: Future-based trajectory prediction for autonomous driving
Qingchao Xu, Yandong Liu 0001, Shixi Wen, Xin Yang 0011
Appl. Intell.4
2025 Efficient spatio-temporal out-of-distribution detection for learning-enabled autonomous systems
Jiankang Ren, Yong Li 0019, Xin Yang 0011
J. Syst. Archit.4
2025 GTIGNet: Global Topology Interaction Graphormer Network for 3D hand pose estimation
Wanshu Fan, Cong Wang 0018, Shixi Wen, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
Neural Networks5
2025 FDC: Feature Dropout Consistency for unsupervised domain adaptation semantic segmentation
Chaoyu Rao, Wanshu Fan, Cong Wang 0018, Xin Yang 0011, Xiaopeng Wei
Neural Networks4
2025 Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
abstract
Humans excel at audiovisual speech recognition (AVSR), motivating the development of human-inspired computing for robust and efficient AVSR models. Spiking neural networks (SNNs), mimicking the brain’s information-processing mechanisms, offer a promising foundation. However, research on SNN-based AVSR remains limited, with most audio-visual methods focusing on object or digit recognition. These methods oversimplify multimodal fusion, neglecting modality-specific characteristics and interactions. Additionally, they often rely on future information, increasing recognition latency and limiting real-time applicability. Inspired by human speech perception, this paper proposes a novel human-inspired SNN named HI-AVSNN for AVSR, incorporating three computing characteristics: spike activity, cueing interaction, and causal processing. For cueing interaction, we introduce a Spike-Driven Visual-Cued Speech Processing (sVCSP) scheme, where visual features hierarchically guide speech processing to enhance critical features. For causal processing, we align the temporal dimensions of SNN with audio-visual inputs and apply temporal masking to ensure only past and current information is used. For spike activity, in addition to SNNs, we incorporate event cameras to capture lip movements as spikes, efficiently encoding visual data like the human retina. Experiments on two event-based AVSR datasets demonstrate our method outperforms existing audio-visual SNN fusion techniques, showcasing the effectiveness, robustness, and efficiency achieved through our human-inspired computing.
Qianhui Liu, Yang Wang 0106, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001
IEEE Trans. Computers4
2025 Spiking Neural Networks With Adaptive Membrane Time Constant for Event-Based Tracking
abstract
The brain-inspired Spiking Neural Networks (SNNs) work in an event-driven manner and have an implicit recurrence in neuronal membrane potential to memorize information over time, which are inherently suitable to handle temporal event-based streams. Despite their temporal nature and recent approaches advancements, these methods have predominantly been assessed on event-based classification tasks. In this paper, we explore the utility of SNNs for event-based tracking tasks. Specifically, we propose a brain-inspired adaptive Leaky Integrate-and-Fire neuron (BA-LIF) that can adaptively adjust the membrane time constant according to the inputs, thereby accelerating the leakage of meaningless noise features and reducing the decay of valuable information. SNNs composed of our proposed BA-LIF neurons can achieve high performance without a careful and time-consuming trial-by-error initialization on the membrane time constant. The adaptive capability of our network is further improved by introducing an extra temporal feature aggregator (TFA) that assigns attention weights over the temporal dimension. Extensive experiments on various event-based tracking datasets validate the effectiveness of our proposed method. We further validate the generalization capability of our method by applying it to other event-classification tasks.
Jiqing Zhang, Malu Zhang, Yuanchen Wang, Qianhui Liu, Haizhou Li 0001, Xin Yang 0011
IEEE Trans. Image Process.7
2024 Exploiting Polarized Material Cues for Robust Car Detection
abstract
Car detection is an important task that serves as a crucial prerequisite for many automated driving functions. The large variations in lighting/weather conditions and vehicle densities of the scenes pose significant challenges to existing car detection algorithms to meet the highly accurate perception demand for safety, due to the unstable/limited color information, which impedes the extraction of meaningful/discriminative features of cars. In this work, we present a novel learning-based car detection method that leverages trichromatic linear polarization as an additional cue to disambiguate such challenging cases. A key observation is that polarization, characteristic of the light wave, can robustly describe intrinsic physical properties of the scene objects in various imaging conditions and is strongly linked to the nature of materials for cars (e.g., metal and glass) and their surrounding environment (e.g., soil and trees), thereby providing reliable and discriminative features for robust car detection in challenging scenes. To exploit polarization cues, we first construct a pixel-aligned RGB-Polarization car detection dataset, which we subsequently employ to train a novel multimodal fusion network. Our car detection network dynamically integrates RGB and polarization features in a request-and-complement manner and can explore the intrinsic material properties of cars across all learning samples. We extensively validate our method and demonstrate that it outperforms state-of-the-art detection methods. Experimental results show that polarization is a powerful cue for car detection. Our code is available at https://github.com/wind1117/AAAI24-PCDNet.
Wen Dong 0008, Haiyang Mei, Ziqi Wei 0001, Ao Jin, Sen Qiu, Qiang Zhang 0008, Xin Yang 0011
AAAI7
2024 Phasic Diversity Optimization for Population-Based Reinforcement Learning
abstract
Reviewing the previous work of diversity Reinforcement Learning, diversity is often obtained via an augmented loss function, which requires a balance between reward and diversity. Generally, diversity optimization algorithms use Multi-armed Bandits algorithms to select the coefficient in the pre-defined space. However, the dynamic distribution of reward signals for MABs or the conflict between quality and diversity limits the performance of these methods. We introduce the Phasic Diversity Optimization (PDO) algorithm, a Population-Based Training framework that separates reward and diversity training into distinct phases instead of optimizing a multi-objective function. In the auxiliary phase, agents with poor performance diversified via determinants will not replace the better agents in the archive. The decoupling of reward and diversity allows us to use an aggressive diversity optimization in the auxiliary phase without performance degradation. Furthermore, we construct a dogfight scenario for aerial agents to demonstrate the practicality of the PDO algorithm. We introduce two implementations of PDO archive and conduct tests in the newly proposed adversarial dogfight and MuJoCo simulations. The results show that our proposed algorithm achieves better performance than baselines.
Jingcheng Jiang, Haiyin Piao, Yihang Hao, Chuanlu Jiang, Ziqi Wei 0001, Xin Yang 0011
ICRA7
2024 Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition
Yang Wang 0106, Haiyang Mei, Qirui Bao, Ziqi Wei 0001, Zheng Shou 0001, Haizhou Li 0001, Bo Dong 0004, Xin Yang 0011
IJCAI8
2024 Prior-Posterior Knowledge Prompting-and-Reasoning for Surgical Visual Question Localized-Answering
abstract
The Surgical Visual Question Localized-Answering (VQLA) aims to locate the specific instance area while responing the associated question, which has the potential to assist junior resident doctors in understanding the surgical process and offering decision support for surgeons. Yet, this task remains a challenging job for data-driven neural networks, due to the serious reliance on surgical scene information provided by posterior knowledge. Hence, we propose a prior-posterior knowledge prompting-and-Reasoning (PPKPR) method to imitate the operational mode of a surgeon. Surgeons systematically inspect each instance within the surgical scene and its associated question, subsequently relying on their accumulated work experience to understand and answer the questions. The PPKPR comprises three modules: prior-posterior multi-domain knowledge prompter (PPMP), prior-posterior instance knowledge prompter (PPIP), and posterior knowledge Reasoner (PKR). Specifically, PPMP aligns prior-posterior multi-domain knowledge, thus prompting model to alleviate the misinterpretations of the textual question. PPIP provides the prior instance knowledge, ensuring model to focus on correct areas in the visual scene. The prompted knowledge is refined by the PKR to reasoning the final answer. Experimental results demonstrate that our method performs favorably against the state-of-the-art methods on the EndoVis-18 and EndoVis-17 datasets.
Peixi Peng, Wanshu Fan, Wenfei Liu, Xin Yang 0011
IJCNN4
2024 Event-intensity Stereo with Cross-modal Fusion and Contrast
abstract
For binocular stereo, traditional cameras excel in capturing fine details and texture information but are limited in terms of dynamic range and their ability to handle rapid motion. On the contrary, event cameras provide pixel-level intensity changes with low latency and a wide dynamic range, albeit at the cost of less detail in their output. It is natural to leverage the strengths of both modalities. We solve this problem by introducing a cross-modal fusion module that learns a visual representation from both sensor inputs. Additionally, we extract and compare dense event-intensity stereo pair features by contrasting “pairs of event-intensity pairs from different views and different modalities and different timestamps”. This provides the flexibility in masking hard negatives and enables networks to effectively combine event-intensity signals within a contrastive learning framework, leading to an improved matching accuracy and facilitating more accurate estimation of disparity. Experimental results validate the effectiveness of our model and the improvement of disparity estimation accuracy.
Shanglai Qu, Tianyu Meng, Haiyin Piao, Xiaopeng Wei, Xin Yang 0011
IROS7
2024 Exploring Matching Rates: From Keypoint Selection to Camera Relocalization
Chengjiang Long, Yifeng Fei, Qianchen Xia, Erwei Yin, Xin Yang 0011
ACM Multimedia7
2024 CSO: Constraint-Guided Space Optimization for Active Scene Mapping
Xuefeng Yin, Chenyang Zhu 0002, Shanglai Qu, Kai Xu 0004, Xin Yang 0011
ACM Multimedia7
2024 Adaptive Vision Transformer for Event-Based Human Pose Estimation
abstract
Event-based human pose estimation has gained popularity due to the benefits of high temporal resolution and high dynamic range offered by event cameras. The inherent spatial sparsity of event data makes discarding less significant regions a straightforward and effective way to decrease the computation. However, implementing this operation in CNNs poses a challenge, as it disrupts the regularity of dense convolutional workload. In this paper, we propose an adaptive vision transformer, a novel efficient backbone for human pose estimation with event cameras. Specifically, we present two adaptive patch and token sampling approaches based on the characteristics of events, thereby reducing the computational load while still achieving comparable performance. Firstly, we design an adaptive patch sampling scheme to eliminate inactivity patches by assessing the entropy of the events before they are inputted into the transformer. Secondly, we further propose an adaptive token reduction strategy to selectively remove less informative tokens in transformer layers through a dynamic token pruning algorithm. To exploit event-based visual cues in human pose estimation tasks, we construct a large-scale frame-event-based dataset, dubbed Event Multi Movement HPE (EventMM HPE). The dataset provides annotation frequencies up to 240 Hz. Extensive experiments demonstrate that our proposed approach outperforms existing state-of-the-art methods in estimation accuracy. The source code and dataset are available at https://github.com/doublemanyu/Adaptive-Vision-Transformer-for-Event-Based-HPE.
Nannan Yu, Jiqing Zhang, Yuji Zhang 0004, Qirui Bao, Xiaopeng Wei, Xin Yang 0011
ACM Multimedia7
2024 A self-supervised anomaly detection algorithm with interpretability
Zhichao Wu 0001, Xin Yang 0011, Xiaopeng Wei, Peijun Yuan, Yuanping Zhang, Jianming Bai
Expert Syst. Appl.2
2024 A Universal Event-Based Plug-In Module for Visual Object Tracking in Degraded Conditions
Jiqing Zhang, Bo Dong 0004, Yingkai Fu, Yuanchen Wang, Xiaopeng Wei, Xin Yang 0011
Int. J. Comput. Vis.7
2024 DSTN: Dynamic Spatio-Temporal Network for Early Fault Warning in Chemical Processes
Chenming Duan, Zhichao Wu 0001, Xirong Xu, Jianmin Zhu, Ziqi Wei 0001, Xin Yang 0011
Knowl. Based Syst.7
2024 Semantic-Aware Detail Search and Feature Constraint for Cross-Resolution Person Re-Identification
abstract
Cross-resolution person re-identification (CRReID) task devotes to identifying the same person from cross-resolution and cross-camera images. Existing CRReID methods learn the identity features of persons by jointly training the super-resolution (SR) and recognition models. These methods achieve sub-optimal results because the design of SR techniques is mostly oriented towards the visual quality of images rather than recognition tasks. To address this deficiency, we propose a Semantic-Aware detail Search and feature Constraint Network (SASC-Net). Specifically, we propose the semantic-aware detail search (SDS) module that is used to customize an SR module by perceiving identity-related semantic information. Then, we devise an Intra-Scale and Inter-Scale Feature Constraint loss function. It ensures that the affinity relationships of the semantic features of repaired images are close to that of high-resolution (HR) images at the scale level, reducing the solution space of the SDS module and promoting the identification module to focus on more discriminative pedestrian features. The effectiveness of our proposed method is validated by the experimental results on five cross-resolution person datasets.
Tiantian Yan, Xin Yang 0011, Qiang Zhang 0008
IEEE Signal Process. Lett.3
2024 Learning Motion-Guided Multi-Scale Memory Features for Video Shadow Detection
abstract
Natural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods.
Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003
IEEE Trans. Circuits Syst. Video Technol.3
2024 Discovering Expert-Level Air Combat Knowledge via Deep Excitatory-Inhibitory Factorized Reinforcement Learning
abstract
Artificial Intelligence (AI) has achieved a wide range of successes in autonomous air combat decision-making recently. Previous research demonstrated that AI-enabled air combat approaches could even acquire beyond human-level capabilities. However, there remains a lack of evidence regarding two major difficulties. First, the existing methods with fixed decision intervals are mostly devoted to solving what to act but merely pay attention to when to act, which occasionally misses optimal decision opportunities. Second, the method of an expert-crafted finite maneuver library leads to a lack of tactics diversity, which is vulnerable to an opponent equipped with new tactics. In view of this, we propose a novel Deep Reinforcement Learning (DRL) and prior knowledge hybrid autonomous air combat tactics discovering algorithm, namely deep E xcitatory-i N hibitory f ACT or I zed maneu VE r ( ENACTIVE ) learning. The algorithm consists of two key modules, i.e., ENHANCE and FACTIVE. Specifically, ENHANCE learns to adjust the air combat decision-making intervals and appropriately seize key opportunities. FACTIVE factorizes maneuvers and then jointly optimizes them with significant tactics diversity increments. Extensive experimental results reveal that the proposed method outperforms state-of-the-art algorithms with a 62% winning rate and further obtains a margin of a 2.85-fold increase in terms of global tactic space coverage. It also demonstrates that a variety of discovered air combat tactics are comparable to human experts’ knowledge.
Haiyin Piao, Shengqi Yang, Hechang Chen, Junnan Li 0008, Xuanqi Peng, Xin Yang 0011, Zhen Yang 0011, Zhixiao Sun, Yi Chang 0001
ACM Trans. Intell. Syst. Technol.7
2024 Eye Gaze Guided Cross-Modal Alignment Network for Radiology Report Generation
abstract
The potential benefits of automatic radiology report generation, such as reducing misdiagnosis rates and enhancing clinical diagnosis efficiency, are significant. However, existing data-driven methods lack essential medical prior knowledge, which hampers their performance. Moreover, establishing global correspondences between radiology images and related reports, while achieving local alignments between images correlated with prior knowledge and text, remains a challenging task. To address these shortcomings, we introduce a novel Eye Gaze Guided Cross-modal Alignment Network (EGGCA-Net) for generating accurate medical reports. Our approach incorporates prior knowledge from radiologists' Eye Gaze Region (EGR) to refine the fidelity and comprehensibility of report generation. Specifically, we design a Dual Fine-Grained Branch (DFGB) and a Multi-Task Branch (MTB) to collaboratively ensure the alignment of visual and textual semantics across multiple levels. To establish fine-grained alignment between EGR-related images and sentences, we introduce the Sentence Fine-grained Prototype Module (SFPM) within DFGB to capture cross-modal information at different levels. Additionally, to learn the alignment of EGR-related image topics, we introduce the Multi-task Feature Fusion Module (MFFM) within MTB to refine the encoder output information. Finally, a specifically designed label matching mechanism is designed to generate reports that are consistent with the anticipated disease states. The experimental outcomes indicate that the introduced methodology surpasses previous advanced techniques, yielding enhanced performance on two extensively used benchmark datasets: Open-i and MIMIC-CXR.
Peixi Peng, Wanshu Fan, Wenfei Liu, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
IEEE J. Biomed. Health Informatics5
2024 MCFNet: Multi-Attentional Class Feature Augmentation Network for Real-Time Scene Parsing
abstract
For real-time scene parsing tasks, capturing multi-scale semantic features and performing effective feature fusion is crucial. However, many existing solutions ignore stripe-shaped things like poles, traffic lights and are so computationally expensive that cannot meet the high real-time requirements. This article presents a novel model, the Multi-Attention Class Feature Augmentation Network (MCFNet) to address this challenge. MCFNet is designed to capture long-range dependencies across different scales with low computational cost and to perform a weighted fusion of feature maps. It features the BAM (Strip Matrix Based Attention Module) for extracting strip objects in images. The BAM module replaces the conventional self-attention method using square matrices with strip matrices, which allows it to focus more on strip objects while reducing computation. Additionally, MCFNet has a parallel branch that focuses on global information based on self-attention to avoid wasting computation. The two branches are merged to enhance the performance of traditional self-attention modules. Experimental results on two mainstream datasets demonstrate the effectiveness of MCFNet. On the Camvid and Cityscapes test sets, MCFNet achieved 207.5 FPS/73.5% mIoU and 136.1 FPS/71.63% mIoU, respectively. The experiments show that MCFNet outperforms other models on the Camvid dataset and can significantly improve the performance of real-time scene parsing tasks.
Xizhong Wang, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Parallel Dense Vision Transformer and Augmentation Network for Occluded Person Re-identification
Chuxia Yang, Wanshu Fan, Ziqi Wei 0001, Xin Yang 0011, Qiang Zhang 0008
CAD/Graphics4
2023 Deep Polarization Reconstruction with PDAVIS Events
abstract
The polarization event camera PDAVIS is a novel bio-inspired neuromorphic vision sensor that reports both conventional polarization frames and asynchronous, continuously per-pixel polarization brightness changes (polarization events) with fast temporal resolution and large dynamic range. A deep neural network method (Polarization FireNet) was previously developed to reconstruct the polarization angle and degree from polarization events for bridging the gap between the polarization event camera and mainstream computer vision. However, Polarization FireNet applies a network pretrained for normal event-based frame reconstruction independently on each of four channels of polarization events from four linear polarization angles, which ignores the correlations between channels and inevitably introduces content inconsistency between the four reconstructed frames, resulting in unsatisfactory polarization reconstruction performance. In this work, we strive to train an effective, yet efficient, DNN model that directly outputs polarization from the input raw polarization events. To this end, we constructed the first large-scale event-to-polarization dataset, which we subsequently employed to train our events-to-polarization network E2P. E2P extracts rich polarization patterns from input polarization events and enhances features through cross-modality context integration. We demonstrate that E2P outperforms Polarization FireNet by a significant margin with no additional computing cost. Experimental results also show that E2P produces more accurate measurement of polarization than the PDAVIS frames in challenging fast and high dynamic range scenes. Code and data are publicly available at: https://github.com/SensorsINI/e2p.
Haiyang Mei, Zuowen Wang, Xin Yang 0011, Xiaopeng Wei, Tobi Delbruck
CVPR3
2023 Frame-Event Alignment and Fusion Network for High Frame Rate Tracking
abstract
Most existing RGB-based trackers target low frame rate benchmarks of around 30 frames per second. This setting restricts the tracker's functionality in the real world, especially for fast motion. Event-based cameras as bioinspired sensors provide considerable potential for high frame rate tracking due to their high temporal resolution. However, event-based cameras cannot offer fine-grained texture information like conventional cameras. This unique complementarity motivates us to combine conventional frames and events for high frame rate object tracking under various challenging conditions. In this paper, we propose an end-to-end network consisting of multi-modality alignment and fusion modules to effectively combine meaningful information from both modalities at different measurement rates. The alignment module is responsible for cross-style and cross-frame-rate alignment between frame and event modalities under the guidance of the moving cues furnished by events. While the fusion module is accountable for emphasizing valuable features and suppressing noise information by the mutual complement between the two modalities. Extensive experiments show that the proposed approach outper-forms state-of-the-art trackers by a significant margin in high frame rate tracking. With the FE240Hz dataset, our approach achieves high frame rate tracking up to 240Hz.
Jiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li 0072, Jinpeng Bai, Xin Yang 0011
CVPR7
2023 Multi-view Spectral Polarization Propagation for Video Glass Segmentation
abstract
In this paper, we present the first polarization-guided video glass segmentation propagation solution (PGVS-Net) that can robustly and coherently propagate glass segmentation in RGB-P video sequences. By leveraging spatiotemporal polarization and color information, our method combines multi-view polarization cues and thus can alleviate the view dependence of single-input intensity variations on glass objects. We demonstrate that our model can outperform glass segmentation on RGB-only video sequences as well as produce more robust segmentation than per-frame RGB-P single-image segmentation methods. To train and validate PGVS-Net, we introduce a novel RGB-P Glass Video dataset (PGV-117) containing 117 video sequences of scenes captured with different types of camera paths, lighting conditions, dynamics, and glass types.
Yu Qiao 0001, Bo Dong 0004, Ao Jin, Seung-Hwan Baek, Felix Heide, Pieter Peers, Xiaopeng Wei, Xin Yang 0011
ICCV9
2023 Single Depth-image 3D Reflection Symmetry and Shape Prediction
abstract
In this paper, we present Iterative Symmetry Completion Network (ISCNet), a single depth-image shape completion method that exploits reflective symmetry cues to obtain more detailed shapes. The efficacy of single depth-image shape completion methods is often sensitive to the accuracy of the symmetry plane. ISCNet therefore jointly estimates the symmetry plane and shape completion iteratively; more complete shapes contribute to more robust symmetry plane estimates and vice versa. Furthermore, our shape completion method operates in the image domain, enabling more efficient high-resolution, detailed geometry reconstruction. We perform the shape completion from pairs of viewpoints, reflected across the symmetry plane, predicted by a reinforcement learning agent to improve robustness and to simultaneously explicitly leverage symmetry. We demonstrate the effectiveness of ISCNet on a variety of object categories on both synthetic and real-scanned datasets.
Zhaoxuan Zhang, Bo Dong 0004, Felix Heide, Pieter Peers, Xin Yang 0011
ICCV7
2023 Spiking Reinforcement Learning for Weakly-Supervised Anomaly Detection
Ao Jin, Zhichao Wu 0001, Qianchen Xia, Xin Yang 0011
ICONIP (5)5
2023 Event-Enhanced Multi-Modal Spiking Neural Network for Dynamic Obstacle Avoidance
abstract
Autonomous obstacle avoidance is of vital importance for an intelligent agent such as a mobile robot to navigate in its environment. Existing state-of-the-art methods train a spiking neural network (SNN) with deep reinforcement learning (DRL) to achieve energy-efficient and fast inference speed in complex/unknown scenes. These methods typically assume that the environment is static while the obstacles in real-world scenes are often dynamic. The movement of obstacles increases the complexity of the environment and poses a great challenge to the existing methods. In this work, we approach robust dynamic obstacle avoidance twofold. First, we introduce the neuromorphic vision sensor (i.e., event camera) to provide motion cues complementary to the traditional Laser depth data for handling dynamic obstacles. Second, we develop an DRL-based event-enhanced multimodal spiking actor network (EEM-SAN) that extracts information from motion events data via unsupervised representation learning and fuses Laser and event camera data with learnable thresholding. Experiments demonstrate that our EEM-SAN outperforms state-of-the-art obstacle avoidance methods by a significant margin, especially for dynamic obstacle avoidance.
Yang Wang 0106, Bo Dong 0004, Yuji Zhang 0004, Yunduo Zhou, Haiyang Mei, Ziqi Wei 0001, Xin Yang 0011
ACM Multimedia7
2023 Camouflaged Object Segmentation with Omni Perception
Haiyang Mei, Ke Xu 0010, Yunduo Zhou, Yang Wang 0106, Haiyin Piao, Xiaopeng Wei, Xin Yang 0011
Int. J. Comput. Vis.7
2023 Large-Field Contextual Feature Learning for Glass Detection
abstract
Glass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind the glass. In this paper, we propose an important problem of detecting glass surfaces from a single RGB image. To address this problem, we construct the first large-scale glass detection dataset (GDD) and propose a novel glass detection network, called GDNet-B, which explores abundant contextual cues in a large field-of-view via a novel large-field contextual feature integration (LCFI) module and integrates both high-level and low-level boundary features with a boundary feature enhancement (BFE) module. Extensive experiments demonstrate that our GDNet-B achieves satisfying glass detection results on the images within and beyond the GDD testing set. We further validate the effectiveness and generalization capability of our proposed GDNet-B by applying it to other vision tasks, including mirror segmentation and salient object detection. Finally, we show the potential applications of glass detection and discuss possible future research directions.
Haiyang Mei, Xin Yang 0011, Letian Yu, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image
abstract
We present a deep reinforcement learning method of progressive view inpainting for colored semantic point cloud scene completion under volume guidance, achieving high-quality scene reconstruction from only a single RGB-D image with severe occlusion. Our approach is end-to-end, consisting of three modules: 3D scene volume reconstruction, 2D RGB-D and segmentation image inpainting, and multi-view selection for completion. Given a single RGB-D image, our method first predicts its semantic segmentation map and goes through the 3D volume branch to obtain a volumetric scene reconstruction as a guide to the next view inpainting step, which attempts to make up the missing information; the third step involves projecting the volume under the same view of the input, concatenating them to complete the current view RGB-D and segmentation map, and integrating all RGB-D and segmentation maps into the point cloud. Since the occluded areas are unavailable, we resort to a A3C network to glance around and pick the next best view for large hole completion progressively until a scene is adequately reconstructed while guaranteeing validity. All steps are learned jointly to achieve robust and consistent results. We perform qualitative and quantitative evaluations with extensive experiments on the 3D-FUTURE data, obtaining better results than state-of-the-arts.
Zhaoxuan Zhang, Xiaoguang Han 0001, Bo Dong 0004, Xin Yang 0011
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Monocular Camera-Based Complex Obstacle Avoidance via Efficient Deep Reinforcement Learning
abstract
Deep reinforcement learning has achieved great success in laser-based collision avoidance works because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to deploy for a large scale of robots but also demonstrate unsatisfactory robustness towards the complex obstacles, including irregular obstacles, e.g., tables, chairs, and shelves, as well as complex ground and special materials. In this paper, we propose a novel monocular camera-based complex obstacle avoidance framework. Particularly, we innovatively transform the captured RGB images to pseudo-laser measurements for efficient deep reinforcement learning. Compared to the traditional laser measurement captured at a certain height that only contains one-dimensional distance information away from the neighboring obstacles, our proposed pseudo-laser measurement fuses the depth and semantic information of the captured RGB image, which makes our method effective for complex obstacles. We also design a feature extraction guidance module to weight the input pseudo-laser measurement, and the agent has more reasonable attention for the current state, which is conducive to improving the accuracy and efficiency of the obstacle avoidance policy. Besides, we adaptively add the synthesized noise to the laser measurement during the training stage to decrease the sim-to-real gap and increase the robustness of our model in the real environment. Finally, the experimental results show that our framework achieves state-of-the-art performance in several virtual and real-world scenarios.
Jianchuan Ding, Lingping Gao, Wenxi Liu, Haiyin Piao, Jia Pan 0001, Zhenjun Du, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.7
2023 Toward Realistic 3D Human Motion Prediction With a Spatio-Temporal Cross- Transformer Approach
abstract
Human motion prediction intends to predict how humans move given a historical sequence of 3D human motions. Recent transformer-based methods have attracted increasing attentions and demonstrated their promising performance in 3D human motion prediction. However, existing methods generally decompose the input of human motion information into spatial and temporal branches in a separate way and seldom consider their inherent coherence between the two branches, hence often failing to register the dynamic spatio-temporal information during the training process. Motivated by these issues, we propose a spatio-temporal cross-transformer network (STCT) for 3D human motion predictions. Specifically, we investigate various types of interaction methods (i.e., Concatenation Interaction, Msg token interaction, and Cross-transformer) to capture the coherence of the spatial and temporal branches. According to the obtained results, the proposed cross-transformer interaction method shows its superiority over other methods. Meanwhile, considering that most existing works treat the human body as a set of 3D human joint positions, the predicted human joints are proportionally less appropriate to the realistic human body due to unreasonable bone length and non-plausible poses as time progresses. We further resort to the bone constraints of human mesh to produce more realistic human motions. By fitting a parametric body model (i.e., SMPL-X model) to the predicted human joints, a reconstruction loss function is proposed to remedy the unreasonable bone length and pose errors. Comprehensive experiments on AMASS and Human3.6M datasets have demonstrated that our method achieves superior performance over compared methods.
Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Wenbin Pei, Hong-Wei Ge, Xin Yang 0011, Qiang Zhang 0008, Mengjie Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 Distractor-Aware Event-Based Tracking
abstract
Event cameras, or dynamic vision sensors, have recently achieved success from fundamental vision tasks to high-level vision researches. Due to its ability to asynchronously capture light intensity changes, event camera has an inherent advantage to capture moving objects in challenging scenarios including objects under low light, high dynamic range, or fast moving objects. Thus event camera are natural for visual object tracking. However, the current event-based trackers derived from RGB trackers simply modify the input images to event frames and still follow conventional tracking pipeline that mainly focus on object texture for target distinction. As a result, the trackers may not be robust dealing with challenging scenarios such as moving cameras and cluttered foreground. In this paper, we propose a distractor-aware event-based tracker that introduces transformer modules into Siamese network architecture (named DANet). Specifically, our model is mainly composed of a motion-aware network and a target-aware network, which simultaneously exploits both motion cues and object contours from event data, so as to discover motion objects and identify the target object by removing dynamic distractors. Our DANet can be trained in an end-to-end manner without any post-processing and can run at over 80 FPS on a single V100. We conduct comprehensive experiments on two large event tracking datasets to validate the proposed model. We demonstrate that our tracker has superior performance against the state-of-the-art trackers in terms of both accuracy and efficiency.
Yingkai Fu, Meng Li 0072, Wenxi Liu, Yuanchen Wang, Jiqing Zhang, Xiaopeng Wei, Xin Yang 0011
IEEE Trans. Image Process.8
2023 S $^3$ Net: Self-Supervised Self-Ensembling Network for Semi-Supervised RGB-D Salient Object Detection
abstract
RGB-D salient object detection aims to detect visually distinctive objects or regions from a pair of the RGB image and the depth image. State-of-the-art RGB-D saliency detectors are mainly based on convolutional neural networks but almost suffer from an intrinsic limitation relying on the labeled data, thus degrading detection accuracy in complex cases. In this work, we present a self-supervised self-ensembling network (S$^3$Net) for semi-supervised RGB-D salient object detection by leveraging the unlabeled data and exploring a self-supervised learning mechanism. To be specific, we first build a self-guided convolutional neural network (SG-CNN) as a baseline model by developing a series of three-layer cross-model feature fusion (TCF) modules to leverage complementary information among depth and RGB modalities and formulating an auxiliary task that predicts a self-supervised image rotation angle. After that, to further explore the knowledge from unlabeled data, we assign SG-CNN to a student network and a teacher network, and encourage the saliency predictions and self-supervised rotation predictions from these two networks to be consistent on the unlabeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our network quantitatively and qualitatively outperforms the state-of-the-art methods.
Lei Zhu 0003, Xiaoqiang Wang 0007, Ping Li 0016, Xin Yang 0011, Qing Zhang 0006, Weiming Wang 0002, Carola-Bibiane Schönlieb, C. L. Philip Chen
IEEE Trans. Multim.4
2023 Mirror Segmentation via Semantic-aware Contextual Contrasted Feature Learning
abstract
Mirrors are everywhere in our daily lives. Existing computer vision systems do not consider mirrors, and hence may get confused by the reflected content inside a mirror, resulting in a severe performance degradation. However, separating the real content outside a mirror from the reflected content inside it is non-trivial. The key challenge is that mirrors typically reflect contents similar to their surroundings, making it very difficult to differentiate the two. In this article, we present a novel method to segment mirrors from a single RGB image. To the best of our knowledge, this is the first work to address the mirror segmentation problem with a computational approach. We make the following contributions: First, we propose a novel network, called MirrorNet+, for mirror segmentation, by modeling both contextual contrasts and semantic associations. Second, we construct the first large-scale mirror segmentation dataset, which consists of 4,018 pairs of images containing mirrors and their corresponding manually annotated mirror masks, covering a variety of daily-life scenes. Third, we conduct extensive experiments to evaluate the proposed method and show that it outperforms the related state-of-the-art detection and segmentation methods. Fourth, we further validate the effectiveness and generalization capability of the proposed semantic awareness contextual contrasted feature learning by applying MirrorNet+ to other vision tasks, i.e., salient object detection and shadow detection. Finally, we provide some applications of mirror segmentation and analyze possible future research directions. Project homepage: https://mhaiyang.github.io/TOMM2022-MirrorNet+/index.html .
Haiyang Mei, Letian Yu, Ke Xu 0010, Yang Wang 0106, Xin Yang 0011, Xiaopeng Wei, Rynson W. H. Lau
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Hierarchical and Progressive Image Matting
abstract
Most matting research resorts to advanced semantics to achieve high-quality alpha mattes, and a direct low-level features combination is usually explored to complement alpha details. However, we argue that appearance-agnostic integration can only provide biased foreground (FG) details and that alpha mattes require different-level feature aggregation for better pixel-wise opacity perception. In this article, we propose an end-to-end hierarchical and progressive attention matting network (HAttMatting++), which can better predict the opacity of the FG from single RGB images without additional input. Specifically, we utilize channel-wise attention (CA) to distill pyramidal features and employ spatial attention (SA) at different levels to filter appearance cues. This progressive attention mechanism can estimate alpha mattes from adaptive semantics and semantics-indicated boundaries. We also introduce a hybrid loss function fusing structural similarity, mean square error, adversarial loss, and sentry supervision to guide the network to further improve the overall FG structure. In addition, we construct a large-scale and challenging image matting dataset comprised of 59,000 training images and 1,000 test images (a total of 646 distinct FG alpha mattes), which can further improve the robustness of our hierarchical and progressive aggregation model. Extensive experiments demonstrate that the proposed HAttMatting++ can capture sophisticated FG structures and achieve state-of-the-art performance with single RGB images as input.
Yu Qiao 0001, Yuhao Liu 0001, Ziqi Wei 0001, Yuxin Wang 0001, Qiang Cai 0001, Guofeng Zhang 0026, Xin Yang 0011
ACM Trans. Multim. Comput. Commun. Appl.7
2023 A Geometrical Approach to Evaluate the Adversarial Robustness of Deep Neural Networks
abstract
Deep neural networks (DNNs) are widely used for computer vision tasks. However, it has been shown that deep models are vulnerable to adversarial attacks—that is, their performances drop when imperceptible perturbations are made to the original inputs, which may further degrade the following visual tasks or introduce new problems such as data and privacy security. Hence, metrics for evaluating the robustness of deep models against adversarial attacks are desired. However, previous metrics are mainly proposed for evaluating the adversarial robustness of shallow networks on the small-scale datasets. Although the Cross Lipschitz Extreme Value for nEtwork Robustness (CLEVER) metric has been proposed for large-scale datasets (e.g., the ImageNet dataset), it is computationally expensive and its performance relies on a tractable number of samples. In this article, we propose the Adversarial Converging Time Score (ACTS), an attack-dependent metric that quantifies the adversarial robustness of a DNN on a specific input. Our key observation is that local neighborhoods on a DNN’s output surface would have different shapes given different inputs. Hence, given different inputs, it requires different time for converging to an adversarial sample. Based on this geometry meaning, the ACTS measures the converging time as an adversarial robustness metric. We validate the effectiveness and generalization of the proposed ACTS metric against different adversarial attacks on the large-scale ImageNet dataset using state-of-the-art deep networks. Extensive experiments show that our ACTS metric is an efficient and effective adversarial metric over the previous CLEVER metric.
Yang Wang 0106, Bo Dong 0004, Ke Xu 0010, Haiyin Piao, Yufei Ding 0001, Xin Yang 0011
ACM Trans. Multim. Comput. Commun. Appl.7
2023 Explore Contextual Information for 3D Scene Graph Generation
abstract
3D scene graph generation (SGG) has been of high interest in computer vision. Although the accuracy of 3D SGG on coarse classification and single relation label has been gradually improved, the performance of existing works is still far from being perfect for fine-grained and multi-label situations. In this article, we propose a framework fully exploring contextual information for the 3D SGG task, which attempts to satisfy the requirements of fine-grained entity class, multiple relation labels, and high accuracy simultaneously. Our proposed approach is composed of a Graph Feature Extraction module and a Graph Contextual Reasoning module, achieving appropriate information-redundancy feature extraction, structured organization, and hierarchical inferring. Our approach achieves superior or competitive performance over previous methods on the 3DSSG dataset, especially on the relationship prediction sub-task.
Chengjiang Long, Zhaoxuan Zhang, Bokai Liu, Qiang Zhang 0008, Xin Yang 0011
IEEE Trans. Vis. Comput. Graph.7
2023 Lightweight single-image super-resolution via multi-scale feature fusion CNN and multiple attention block
Wanshu Fan, Xin Yang 0011, Qiang Zhang 0008
Vis. Comput.3
2022 CPRAL: Collaborative Panoptic-Regional Active Learning for Semantic Segmentation
abstract
Acquiring the most representative examples via active learning (AL) can benefit many data-dependent computer vision tasks by minimizing efforts of image-level or pixel-wise annotations. In this paper, we propose a novel Collaborative Panoptic-Regional Active Learning framework (CPRAL) to address the semantic segmentation task. For a small batch of images initially sampled with pixel-wise annotations, we employ panoptic information to initially select unlabeled samples. Considering the class imbalance in the segmentation dataset, we import a Regional Gaussian Attention module (RGA) to achieve semantics-biased selection. The subset is highlighted by vote entropy and then attended by Gaussian kernels to maximize the biased regions. We also propose a Contextual Labels Extension (CLE) to boost regional annotations with contextual attention guidance. With the collaboration of semantics-agnostic panoptic matching and region-biased selection and extension, our CPRAL can strike a balance between labeling efforts and performance and compromise the semantics distribution. We perform extensive experiments on Cityscapes and BDD10K datasets and show that CPRAL outperforms the cutting-edge methods with impressive results and less labeling proportion.
Yu Qiao 0001, Jincheng Zhu, Chengjiang Long, Zeyao Zhang, Yuxin Wang 0001, Zhenjun Du, Xin Yang 0011
AAAI7
2022 Wider and Higher: Intensive Integration and Global Foreground Perception for Image Matting
Yu Qiao 0001, Ziqi Wei 0001, Yuhao Liu 0001, Yuxin Wang 0001, Qiang Zhang 0008, Xin Yang 0011
CGI7
2022 Glass Segmentation using Intensity and Spectral Polarization Cues
abstract
Transparent and semi-transparent materials pose significant challenges for existing scene understanding and segmentation algorithms due to their lack of RGB texture which impedes the extraction of meaningful features. In this work, we exploit that the light-matter interactions on glass materials provide unique intensity-polarization cues for each observed wavelength of light. We present a novel learning-based glass segmentation network that leverages both trichromatic (RGB) intensities as well as trichromatic linear polarization cues from a single photograph captured without making any assumption on the polarization state of the illumination. Our novel network architecture dynamically fuses and weights both the trichromatic color and polarization cues using a novel global-guidance and multi-scale self-attention module, and leverages global cross-domain contextual information to achieve robust segmentation. We train and extensively validate our segmentation method on a new large-scale RGB-Polarization dataset (RGBP-Glass), and demonstrate that our method outperforms state-of-the-art segmentation approaches by a significant margin.
Haiyang Mei, Bo Dong 0004, Wen Dong 0008, Seung-Hwan Baek, Felix Heide, Pieter Peers, Xiaopeng Wei, Xin Yang 0011
CVPR9
2022 Bi-directional Object-Context Prioritization Learning for Saliency Ranking
abstract
The saliency ranking task is recently proposed to study the visual behavior that humans would typically shift their attention over different objects of a scene based on their degrees of saliency. Existing approaches focus on learning either object-object or object-scene relations. Such a strategy follows the idea of object-based attention in Psychology, but it tends to favor objects with strong semantics (e.g., humans), resulting in unrealistic saliency ranking. We observe that spatial attention works concurrently with object-based attention in the human visual recognition system. During the recognition process, the human spatial attention mechanism would move, engage, and disengage from region to region (i.e., context to context). This inspires us to model region-level interactions, in addition to object-level reasoning, for saliency ranking. Hence, we propose a novel bi-directional method to unify spatial attention and object-based attention for saliency ranking. Our model has two novel modules: (1) a selective object saliency (SOS) module to model object-based attention via inferring the semantic representation of salient objects, and (2) an object-context-object relation (OCOR) module to allocate saliency ranks to objects by jointly modeling object-context and context-object interactions of salient objects. Extensive experiments show that our approach outperforms existing state-of-the-art methods. Code and pretrained model are available at https://github.com/GrassBro/OCOR.
Xin Tian 0015, Ke Xu 0010, Xin Yang 0011, Rynson W. H. Lau
CVPR3
2022 Spiking Transformers for Event-based Single Object Tracking
abstract
Event-based cameras bring a unique capability to tracking, being able to function in challenging real-world conditions as a direct result of their high temporal resolution and high dynamic range. These imagers capture events asynchronously that encode rich temporal and spatial information. However, effectively extracting this information from events remains an open challenge. In this work, we propose a spiking transformer network, STNet, for single object tracking. STNet dynamically extracts and fuses information from both temporal and spatial domains. In particular, the proposed architecture features a transformer module to provide global spatial information and a spiking neural network (SNN) module for extracting temporal cues. The spiking threshold of the SNN module is dynamically adjusted based on the statistical cues of the spatial information, which we find essential in providing robust SNN features. We fuse both feature branches dynamically with a novel cross-domain attention fusion algorithm. Extensive experiments on three event-based datasets, FE240hz, EED and VisEvent validate that the proposed STNet outperforms existing state-of-the-art methods in both tracking accuracy and speed with a significant margin. The code and pretrained models are at https://github.com/Jee-King/CVPR2022_STNet.
Jiqing Zhang, Bo Dong 0004, Haiwei Zhang 0003, Jianchuan Ding, Felix Heide, Xin Yang 0011
CVPR7
2022 Biologically Inspired Dynamic Thresholds for Spiking Neural Networks
abstract
The dynamic membrane potential threshold, as one of the essential properties of a biological neuron, is a spontaneous regulation mechanism that maintains neuronal homeostasis, i.e., the constant overall spiking firing rate of a neuron. As such, the neuron firing rate is regulated by a dynamic spiking threshold, which has been extensively studied in biology. Existing work in the machine learning community does not employ bioinspired spiking threshold schemes. This work aims at bridging this gap by introducing a novel bioinspired dynamic energy-temporal threshold (BDETT) scheme for spiking neural networks (SNNs). The proposed BDETT scheme mirrors two bioplausible observations: a dynamic threshold has 1) a positive correlation with the average membrane potential and 2) a negative correlation with the preceding rate of depolarization. We validate the effectiveness of the proposed BDETT on robot obstacle avoidance and continuous control tasks under both normal conditions and various degraded conditions, including noisy observations, weights, and dynamic environments. We find that the BDETT outperforms existing static and heuristic threshold approaches by significant margins in all tested conditions, and we confirm that the proposed bioinspired dynamic threshold scheme offers homeostasis to SNNs in complex real-world tasks.
Jianchuan Ding, Bo Dong 0004, Felix Heide, Yufei Ding 0001, Yunduo Zhou, Xin Yang 0011
NeurIPS7
2022 GlassNet: Label Decoupling-based Three-stream Neural Network for Robust Image Glass Detection
abstract
Abstract Most of the existing object detection methods generate poor glass detection results, due to the fact that the transparent glass shares the same appearance with arbitrary objects behind it in an image. Different from traditional deep learning‐based wisdoms that simply use the object boundary as an auxiliary supervision, we exploit label decoupling to decompose the original labelled ground‐truth (GT) map into an interior‐diffusion map and a boundary‐diffusion map. The GT map in collaboration with the two newly generated maps breaks the imbalanced distribution of the object boundary, leading to improved glass detection quality. We have three key contributions to solve the transparent glass detection problem: (1) We propose a three‐stream neural network (call GlassNet for short) to fully absorb beneficial features in the three maps. (2) We design a multi‐scale interactive dilation module to explore a wider range of contextual information. (3) We develop an attention‐based boundary‐aware feature Mosaic module to integrate multi‐modal information. Extensive experiments on the benchmark dataset exhibit clear improvements of our method over SOTAs, in terms of both the overall glass detection accuracy and boundary clearness.
Ding Shi, Xuefeng Yan 0001, Dong Liang 0008, Mingqiang Wei, Xin Yang 0011, Yanwen Guo 0001, Haoran Xie 0001
Comput. Graph. Forum6
2022 Adaptive feedback connection with a single-level feature for object detection
abstract
Abstract From the perspective of detector optimisation, detecting objects using only a one‐level feature cannot provide good performance for a wide range of scales. Various complex feature pyramidal structures address this problem using the divide‐and‐conquer strategy and multi‐scale feature fusion. However, this requires adding too many additional convolutional layers and fusion operations. To address the issue, a simple detection part is proposed, which includes three components, namely a one‐level feature map for detection, the encoder structure with feedback connection, and a decoupled head. The redesigned encoder and decoupled head can successfully address the performance decline caused by the one‐level feature‐based detection. Moreover, the proposed method can accelerate the convergence of the detector and achieve a faster inference time. Based on the optimised detection part, an adaptive feedback connection with a single‐level feature (AFS) is proposed for object detection. The experiments conducted on the MS COCO 2017 benchmark show that the proposed method can achieve comparable results with its multi‐scale pyramid counterpart, You Only Look Once v4 (YOLOv4). In addition, AFS can help the YOLOv4 achieve 44.9 mAP at 27 frame per second and converging 82 epochs earlier under the image size of 608×608, which represents a 42.1% improvements in the convergence speed.
Zhongling Ruan, Jianzhong Cao, Huinan Guo, Xin Yang 0011
IET Comput. Vis.5
2022 Learning to Detect Instance-Level Salient Objects Using Complementary Image Labels
Xin Tian 0015, Ke Xu 0010, Xin Yang 0011, Rynson W. H. Lau
Int. J. Comput. Vis.3
2022 Perception-oriented Single Image Super-Resolution Network with Receptive Field Block
Yaqing Hou, Wanshu Fan, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
Neural Comput. Appl.4
2022 Salient object detection with image-level binary supervision
Pengjie Wang 0001, Ying Cao 0001, Xin Yang 0011, Huchuan Lu, Rynson W. H. Lau
Pattern Recognit.4
2022 Exploring Dense Context for Salient Object Detection
abstract
Contexts play an important role in salient object detection (SOD). High-level contexts describe the relations between different parts/objects and thus are helpful for discovering the specific locations of salient objects while low-level contexts could provide the fine detail information for delineating the boundary of the salient objects. However, the way of perceiving/leveraging rich contexts has not been fully investigated by existing SOD works. The common context extraction strategies (e.g., leveraging convolutions with large kernels or atrous convolutions with large dilation rates) do not consider the effectiveness and efficiency simultaneously and may cause sub-optimal solutions. In this paper, we devote to exploring an effective and efficient way to learn rich contexts for accurate SOD. Specifically, we first build a dense context exploration (DCE) module to capture dense multi-scale contexts and further leverage the learned contexts to enhance the features discriminability. Then, we embed multiple DCE modules in an encoder-decoder architecture to harvest dense contexts of different levels. Furthermore, we propose an attentive skip-connection to transmit useful features from the encoder part to the decoder part for better dense context exploration. Finally, extensive experiments demonstrate that the proposed method achieves more superior detection results on the six benchmark datasets than 18 state-of-the-art SOD methods.
Haiyang Mei, Ziqi Wei 0001, Xiaopeng Wei, Qiang Zhang 0008, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.7
2022 A Two-Stage Attentive Network for Single Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and contribute remarkable progress. However, most of the existing CNNs-based SISR methods do not adequately explore contextual information in the feature extraction stage and pay little attention to the final high-resolution (HR) image reconstruction step, hence hindering the desired SR performance. To address the above two issues, in this paper, we propose a two-stage attentive network (TSAN) for accurate SISR in a coarse-to-fine manner. Specifically, we design a novel multi-context attentive block (MCAB) to make the network focus on more informative contextual features. Moreover, we present an essential refined attention block (RAB) which could explore useful cues in HR space for reconstructing fine-detailed HR image. Extensive evaluations on four benchmark datasets demonstrate the efficacy of our proposed TSAN in terms of quantitative metrics and visual effects. Code is available athttps://github.com/Jee-King/TSAN.
Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Haiyin Piao, Haiyang Mei, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.6
2022 Progressive Glass Segmentation
abstract
Glass is very common in the real world. Influenced by the uncertainty about the glass region and the varying complex scenes behind the glass, the existence of glass poses severe challenges to many computer vision tasks, making glass segmentation as an important computer vision task. Glass does not have its own visual appearances but only transmit/reflect the appearances of its surroundings, making it fundamentally different from other common objects. To address such a challenging task, existing methods typically explore and combine useful cues from different levels of features in the deep network. As there exists a characteristic gap between level-different features, i.e., deep layer features embed more high-level semantics and are better at locating the target objects while shallow layer features have larger spatial sizes and keep richer and more detailed low-level information, fusing these features naively thus would lead to a sub-optimal solution. In this paper, we approach the effective features fusion towards accurate glass segmentation in two steps. First, we attempt to bridge the characteristic gap between different levels of features by developing a Discriminability Enhancement (DE) module which enables level-specific features to be a more discriminative representation, alleviating the features incompatibility for fusion. Second, we design a Focus-and-Exploration Based Fusion (FEBF) module to richly excavate useful information in the fusion process by highlighting the common and exploring the difference between level-different features. Combining these two steps, we construct a Progressive Glass Segmentation Network (PGSNet) which uses multiple DE and FEBF modules to progressively aggregate features from high-level to low-level, implementing a coarse-to-fine glass segmentation. In addition, we build the first home-scene-oriented glass segmentation dataset for advancing household robot applications and in-depth research on this topic. Extensive experiments demonstrate that our method outperforms 26 cutting-edge models on three challenging datasets under four standard metrics. The code and dataset will be made publicly available.
Letian Yu, Haiyang Mei, Wen Dong 0008, Ziqi Wei 0001, Yuxin Wang 0001, Xin Yang 0011
IEEE Trans. Image Process.7
2022 Learning Smooth Motion Planning for Intelligent Aerial Transportation Vehicles by Stable Auxiliary Gradient
abstract
Deep Reinforcement Learning (DRL) has been widely attempted for solving real-time intelligent aerial transportation vehicle motion planning tasks recently. When interacting with environment, DRL-driven aerial vehicles inevitably switch the steering actions in high frequency during both exploration and execution phase, resulting in the well known flight trajectory oscillation issue, which makes flight dynamics unstable, and even endangers flight safety in serious cases. Unfortunately, there is hardly any literature about achieving flight trajectory smoothness in DRL-based motion planning. In view of this, we originally formalize the practical flight trajectory smoothen problem as a three-level Nested pArameterized Smooth Trajectory Optimization (NASTO) form. On this basis, a novel Stable Auxiliary Gradient (SAG) algorithm is proposed, which significantly smoothens the DRL-generated flight motions by constructing two independent optimization aspects: the major gradient, and the stable auxiliary gradient. Experimental result reveals that the proposed SAG algorithm outperforms baseline DRL-based intelligent aerial transportation vehicle motion planning algorithms in terms of both learning efficiency and flight motion smoothness.
Haiyin Piao, Li Mo 0001, Xin Yang 0011, Zhixiao Sun, Zhen Yang 0011
IEEE Trans. Intell. Transp. Syst.4
2022 Prior-Induced Information Alignment for Image Matting
abstract
Image matting is an ill-posed problem that aims to estimate the opacity of foreground pixels in an image. However, most existing deep learning-based methods still suffer from the coarse-grained details. In general, these algorithms are incapable of felicitously distinguishing the degree of exploration between deterministic domains (e.g.certain FG and BG pixels) and undetermined domains (e.g.uncertain in-between pixels), or inevitably lose information in the continuous sampling process, leading to a sub-optimal result. In this paper, we propose a novel network named Prior-Induced Information Alignment Matting Network (PIIAMatting), which can efficiently model the distinction of pixel-wise response maps and the correlation of layer-wise feature maps. It mainly consists of a Dynamic Gaussian Modulation mechanism (DGM) and an Information Alignment strategy (IA). Specifically, the DGM can dynamically acquire a pixel-wise domain response map learned from the prior distribution. The response map can present the relationship between the opacity variation and the convergence process during training. On the other hand, the IA comprises an Information Match Module (IMM) and an Information Aggregation Module (IAM), jointly scheduled to match and aggregate the adjacent layer-wise features adaptively. Besides, we also develop a Multi-Scale Refinement (MSR) module to integrate multi-scale receptive field information at the refinement stage to recover the fluctuating appearance details. Extensive quantitative and qualitative evaluations demonstrate that the proposed PIIAMatting performs favourably against state-of-the-art image matting methods on theAlphamatting.com, Composition-1 K and Distinctions-646 dataset.
Yuhao Liu 0001, Jiake Xie, Yu Qiao 0001, Xin Yang 0011
IEEE Trans. Multim.5
2022 Easy2Hard: Learning to Solve the Intractables From a Synthetic Dataset for Structure-Preserving Image Smoothing
abstract
Image smoothing is a prerequisite for many computer vision and graphics applications. In this article, we raise an intriguing question whether a dataset that semantically describes meaningful structures and unimportant details can facilitate a deep learning model to smooth complex natural images. To answer it, we generate ground-truth labels from easy samples by candidate generation and a screening test and synthesize hard samples in structure-preserving smoothing by blending intricate and multifarious details with the labels. To take full advantage of this dataset, we present a joint edge detection and structure-preserving image smoothing neural network (JESS-Net). Moreover, we propose the distinctive total variation loss as prior knowledge to narrow the gap between synthetic and real data. Experiments on different datasets and real images show clear improvements of our method over the state of the arts in terms of both the image cleanness and structure-preserving ability. Code and dataset are available at https://github.com/YidFeng/Easy2Hard.
Yidan Feng, Xuefeng Yan 0001, Xin Yang 0011, Mingqiang Wei, Ligang Liu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Depth-Aware Mirror Segmentation
abstract
We present a novel mirror segmentation method that leverages depth estimates from ToF-based cameras as an additional cue to disambiguate challenging cases where the contrast or relation in RGB colors between the mirror reflection and the surrounding scene is subtle. A key observation is that ToF depth estimates do not report the true depth of the mirror surface, but instead return the total length of the reflected light paths, thereby creating obvious depth dis-continuities at the mirror boundaries. To exploit depth information in mirror segmentation, we first construct a large-scale RGB-D mirror segmentation dataset, which we subsequently employ to train a novel depth-aware mirror segmentation framework. Our mirror segmentation framework first locates the mirrors based on color and depth discontinuities and correlations. Next, our model further refines the mirror boundaries through contextual contrast taking into account both color and depth information. We extensively validate our depth-aware mirror segmentation method and demonstrate that our model outperforms state-of-the-art RGB and RGB-D based methods for mirror segmentation. Experimental results also show that depth is a powerful cue for mirror segmentation.
Haiyang Mei, Bo Dong 0004, Wen Dong 0008, Pieter Peers, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
CVPR5
2021 Camouflaged Object Segmentation With Distraction Mining
abstract
Camouflaged object segmentation (COS) aims to identify objects that are "perfectly" assimilate into their surroundings, which has a wide range of valuable applications. The key challenge of COS is that there exist high intrinsic similarities between the candidate objects and noise background. In this paper, we strive to embrace challenges towards effective and efficient COS. To this end, we develop a bio-inspired framework, termed Positioning and Focus Network (PFNet), which mimics the process of predation in nature. Specifically, our PFNet contains two key modules, i.e., the positioning module (PM) and the focus module (FM). The PM is designed to mimic the detection process in predation for positioning the potential target objects from a global perspective and the FM is then used to perform the identification process in predation for progressively refining the coarse prediction via focusing on the ambiguous regions. Notably, in the FM, we develop a novel distraction mining strategy for the distraction discovery and removal, to benefit the performance of estimation. Extensive experiments demonstrate that our PFNet runs in real-time (72 FPS) and significantly outperforms 18 cutting-edge models on three challenging datasets under four standard metrics.
Haiyang Mei, Ge-Peng Ji, Ziqi Wei 0001, Xin Yang 0011, Xiaopeng Wei, Deng-Ping Fan
CVPR4
2021 Tripartite Information Mining and Integration for Image Matting
abstract
With the development of deep convolutional neural networks, image matting has ushered in a new phase. Regarding the nature of image matting, most researches have focused on solutions for transition regions. However, we argue that many existing approaches are excessively focused on transition-dominant local fields and ignored the inherent coordination between global information and transition optimisation. In this paper, we propose the Tripartite Information Mining and Integration Network (TIMI-Net) to harmonize the coordination between global and local attributes formally. Specifically, we resort to a novel 3-branch encoder to accomplish comprehensive mining of the input information, which can supplement the neglected coordination between global and local fields. In order to achieve effective and complete interaction between such multi-branches information, we develop the Tripartite Information Integration (T I2) Module to transform and integrate the interconnections between the different branches. In addition, we built a large-scale human matting dataset (Human-2K) to advance human image matting, which consists of 2100 high-precision human images (2000 images for training and 100 images for test). Finally, we con-duct extensive experiments to prove the performance of our proposed TIMI-Net, which demonstrates that our method performs favourably against the SOTA approaches on the alphamatting.com (Rank First), Composition-1K (MSE- 0.006, Grad-11.5), Distinctions-646 and our Human-2K. Also, we have developed an online evaluation website to perform natural image matting.
Yuhao Liu 0001, Jiake Xie, Xiao Shi 0004, Yu Qiao 0001, Xin Yang 0011
ICCV7
2021 Object Tracking by Jointly Exploiting Frame and Event Domain
abstract
Inspired by the complementarity between conventional frame-based and bio-inspired event-based cameras, we propose a multi-modal based approach to fuse visual cues from the frame- and event-domain to enhance the single object tracking performance, especially in degraded conditions (e.g., scenes with high dynamic range, low light, and fast-motion objects). The proposed approach can effectively and adaptively combine meaningful information from both domains. Our approach’s effectiveness is enforced by a novel designed cross-domain attention schemes, which can effectively enhance features based on self- and cross-domain attention schemes; The adaptiveness is guarded by a specially designed weighting scheme, which can adaptively balance the contribution of the two domains. To exploit event-based visual cues in single-object tracking, we construct a large-scale frame-event-based dataset, which we subsequently employ to train a novel frame-event fusion based model. Extensive experiments show that the proposed approach outperforms state-of-the-art frame-based tracking methods by at least 10.4% and 11.9% in terms of representative success rate and precision rate, respectively. Besides, the effectiveness of each key component of our approach is evidenced by our thorough ablation study.
Jiqing Zhang, Xin Yang 0011, Yingkai Fu, Xiaopeng Wei, Bo Dong 0004
ICCV2
2021 A Vision-based Irregular Obstacle Avoidance Framework via Deep Reinforcement Learning
abstract
Deep reinforcement learning has achieved great success in laser-based collision avoidance work because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to apply on a large scale but also have poor robustness to irregular objects, e.g., tables, chairs, shelves, etc. In this paper, we propose a vision-based collision avoidance framework to solve the challenging problem. Our method attempts to estimate the depth and incorporate the semantic information from RGB data to obtain a new form of data, pseudo-laser data, which combines the advantages of visual information and laser information. Compared to traditional laser data that only contains the one-dimensional distance information captured at a certain height, our proposed pseudo-laser data encodes the depth information and semantic information within the image, which makes our method more effective for irregular obstacles. Besides, we adaptively add noise to the laser data during the training stage to increase the robustness of our model in the real world, due to the estimated depth information is not accurate. Experimental results show that our framework achieves state-of-the-art performance in several unseen virtual and real-world scenarios.
Lingping Gao, Jianchuan Ding, Wenxi Liu, Haiyin Piao, Yuxin Wang 0001, Xin Yang 0011
IROS6
2021 MBKD: Acceleration structure designed for moving primitives
Haiyin Piao, Pengyuan Du, Letian Yu, Yuxin Wang 0001, Xin Yang 0011
Comput. Graph.7
2021 ASFNet: Adaptive multiscale segmentation fusion network for real-time semantic segmentation
abstract
Abstract Recently, the development of deep learning has facilitated continuous progress in the field of computer vision. Pixel‐level semantic segmentation serves as a fundamental task in computer vision. It achieves significant results by connecting wider and deeper backbone networks and building fine‐grained segmentation heads. However, applications such as self‐driving cars are more critical to the computational speed of the algorithms. The trade‐off between accuracy and real‐time performance of existing algorithms is still a challenging task. To address this challenge, this article proposes an adaptive multiscale segmentation fusion network to fuse multiscale contextual, which designs an adaptive multiscale segmentation fusion module based on an attention mechanism. Using segmentation fusion instead of feature fusion, the multiscale segmentation results are aggregated to obtain more precise segmentation results. The final results achieved 70.9% mIoU of accuracy in the Cityspace test set, processing images at 61 FPS when the input is 1024 × 2048. In addition, when adjusting the input size to 512 × 1024, the images are processed at 185 FPS.
Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
Comput. Animat. Virtual Worlds3
2021 Intensity-Aware Single-Image Deraining With Semantic and Color Regularization
abstract
Rain degrades image visual quality and disrupts object structures, obscuring their details and erasing their colors. Existing deraining methods are primarily based on modeling either visual appearances of rain or its physical characteristics (e.g., rain direction and density), and thus suffer from two common problems. First, due to the stochastic nature of rain, they tend to fail in recognizing rain streaks correctly, and wrongly remove image structures and details. Second, they fail to recover the image colors erased by heavy rain. In this paper, we address these two problems with the following three contributions. First, we propose a novel PHP block to aggregate comprehensive spatial and hierarchical information for removing rain streaks of different sizes. Second, we propose a novel network to first remove rain streaks, then recover objects structures/colors, and finally enhance details. Third, to train the network, we prepare a new dataset, and propose a novel loss function to introduce semantic and color regularization for deraining. Extensive experiments demonstrate the superiority of the proposed method over state-of-the-art deraining methods on both synthesized and real-world data, in terms of visual quality, quantitative accuracy, and running speed.
Ke Xu 0010, Xin Tian 0015, Xin Yang 0011, Rynson W. H. Lau
IEEE Trans. Image Process.3
2021 Automatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon Generation
abstract
In this article, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes by analyzing the subtitles and stylizes keyframes into comic-style images. Then, we propose a novel automatic multi-page layout framework that can allocate the images across multiple pages and synthesize visually interesting layouts based on the rich semantics of the images (e.g., importance and inter-image relation). Finally, as opposed to using the same type of balloon as in previous works, we propose an emotion-aware balloon generation method to create different types of word balloons by analyzing the emotion of subtitles and audio. Our method is able to vary balloon shapes and word sizes in balloons in response to different emotions, leading to more enriched reading experience. Once the balloons are generated, they are placed adjacent to their corresponding speakers via speaker detection. Our results show that our method, without requiring any user inputs, can generate high-quality comic pages with visually rich layouts and balloons. Our user studies also demonstrate that users prefer our generated results over those by state-of-the-art comic generation systems.
Xin Yang 0011, Zongliang Ma, Letian Yu, Ying Cao 0001, Xiaopeng Wei, Qiang Zhang 0008, Rynson W. H. Lau
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Smart Scribbles for Image Matting
abstract
Image matting is an ill-posed problem that usually requires additional user input, such as trimaps or scribbles. Drawing a fine trimap requires a large amount of user effort, while using scribbles can hardly obtain satisfactory alpha mattes for non-professional users. Some recent deep learning–based matting networks rely on large-scale composite datasets for training to improve performance, resulting in the occasional appearance of obvious artifacts when processing natural images. In this article, we explore the intrinsic relationship between user input and alpha mattes and strike a balance between user effort and the quality of alpha mattes. In particular, we propose an interactive framework, referred to as smart scribbles, to guide users to draw few scribbles on the input images to produce high-quality alpha mattes. It first infers the most informative regions of an image for drawing scribbles to indicate different categories (foreground, background, or unknown) and then spreads these scribbles (i.e., the category labels) to the rest of the image via our well-designed two-phase propagation. Both neighboring low-level affinities and high-level semantic features are considered during the propagation process. Our method can be optimized without large-scale matting datasets and exhibits more universality in real situations. Extensive experiments demonstrate that smart scribbles can produce more accurate alpha mattes with reduced additional input, compared to the state-of-the-art matting methods.
Xin Yang 0011, Yu Qiao 0001, Shaozhe Chen, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Multi-domain collaborative feature representation for robust visual object tracking
Jiqing Zhang, Bo Dong 0004, Yingkai Fu, Yuxin Wang 0001, Xin Yang 0011
Vis. Comput.6
2020 Reweighted Non-convex Non-smooth Rank Minimization Based Spectral Clustering on Grassmann Manifold
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011
ACCV (5)5
2020 Efficient Attention Calibration Network for Real-Time Semantic Segmentation
abstract
In recent years, the attention mechanism has been widely used in computer vision. Semantic segmentation, as one of the fundamental tasks of computer vision, has been subject to tremendous development as a result. But because of its huge computing overhead, attention-based approaches are difficult to use for real-time applications such as self-driving. In this paper, we propose a self-calibration method baesd on self-attentiion that successfully applies the attention mechanism to real-time semantic segmentation. Specifically, a spatial attention module to adjust the edges of the coarse segmentation results which gained from the real-time semantic segmentation backbone network, and obtain more granular segmentation results. We refer to this method as the Efficient Attentional Calibration Network (EACNet). Experiments on the Cityscapes dataset validate the efficiency and performance of the method. With the high-resolution input and without any post-processing, EACNet achieved 72.4% mIoU of accuracy while running at 116.9 FPS. Compared to other state-of-the-art methods for real-time semantic segmentation, our network gained a better balance between performance and speed.
Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
ACML4
2020 Weakly-supervised Salient Instance Detection
Xin Tian 0015, Ke Xu 0010, Xin Yang 0011, Rynson W. H. Lau
BMVC3
2020 Memetic Multi-agent optimization with Problem Reformulation by Coordinate Rotation
abstract
Memetic multi-agent system (MeMAS) is recently proposed as an enhanced version that integrates meme concept into multi-agent system (MAS) wherein all meme-inspired agents have an improvement in learning performance via meme evolution independently or social interaction. In the process of solving the black box optimization problem, the potential advantages of MeMAS have not been utilized well, which makes it a fertile area for further exploration. This paper presents a memetic multi-agent optimization paradigm through coordinate rotation (MeMAO-R) to combine MeMAS with evolutionary algorithms (EAs) to improve optimization efficiency. Based on MeMAS, the particular interest of MeMAO-R is placed on assisting original complex optimization task with new tasks generated by coordinate rotation. Further, MeMAO-R constructs the social interaction mechanism which facilitates to improve their convergence speed for solving the target optimization problem by utilizing meaningful information transferred across multiple agents with differing views of the target problem. Besides, MeMAO-R employs one or more classical EAs as the fundamental population based evolutionary solvers for multiple agents to optimize multiple tasks in a multi-agent scenario. Lastly, to testify the efficacy of the proposed MeMAO-R, comprehensive empirical studies on basic optimization problems are provided.
Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Abhishek Gupta 0001, Xianneng Li
CEC5
2020 Learning to Restore Low-Light Images via Decomposition-and-Enhancement
abstract
Low-light images typically suffer from two problems. First, they have low visibility (i.e., small pixel values). Second, noise becomes significant and disrupts the image content, due to low signal-to-noise ratio. Most existing low-light image enhancement methods, however, learn from noise-negligible datasets. They rely on users having good photographic skills in taking images with low noise. Unfortunately, this is not the case for majority of the low-light images. While concurrently enhancing a low-light image and removing its noise is ill-posed, we observe that noise exhibits different levels of contrast in different frequency layers, and it is much easier to detect noise in the low-frequency layer than in the high one. Inspired by this observation, we propose a frequency-based decomposition- and- enhancement model for low-light image enhancement. Based on this model, we present a novel network that first learns to recover image objects in the low-frequency layer and then enhances high-frequency details based on the recovered image objects. In addition, we have prepared a new low-light image dataset with real noise to facilitate learning. Finally, we have conducted extensive experiments to show that the proposed method outperforms state-of-the-art approaches in enhancing practical noisy low-light images.
Ke Xu 0010, Xin Yang 0011, Rynson W. H. Lau
CVPR2
2020 Attention Scaling for Crowd Counting
abstract
Convolutional Neural Network (CNN) based methods generally take crowd counting as a regression task by outputting crowd densities. They learn the mapping between image contents and crowd density distributions. Though having achieved promising results, these data-driven counting networks are prone to overestimate or underestimate people counts of regions with different density patterns, which degrades the whole count accuracy. To overcome this problem, we propose an approach to alleviate the counting performance differences in different regions. Specifically, our approach consists of two networks named Density Attention Network (DANet) and Attention Scaling Network (ASNet). DANet provides ASNet with attention masks related to regions of different density levels. ASNet first generates density maps and scaling factors and then multiplies them by attention masks to output separate attention-based density maps. These density maps are summed to give the final density map. The attention scaling factors help attenuate the estimation errors in different regions. Furthermore, we present a novel Adaptive Pyramid Loss (APLoss) to hierarchically calculate the estimation losses of sub-regions, which alleviates the training bias. Extensive experiments on four challenging datasets (ShanghaiTech Part A, UCF_CC_50, UCF-QNRF, and WorldExpo'10) demonstrate the superiority of the proposed approach.
Xiaoheng Jiang, Li Zhang 0072, Mingliang Xu 0001, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Xin Yang 0011, Yanwei Pang
CVPR7
2020 Don't Hit Me! Glass Detection in Real-World Scenes
abstract
Glass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind the glass, and the content within the glass region is typically similar to those behind it. In this paper, we propose an important problem of detecting glass from a single RGB image. To address this problem, we construct a large-scale glass detection dataset (GDD) and design a glass detection network, called GDNet, which explores abundant contextual cues for robust glass detection with a novel large-field contextual feature integration (LCFI) module. Extensive experiments demonstrate that the proposed method achieves more superior glass detection results on our GDD test set than state-of-the-art methods fine-tuned for glass detection.
Haiyang Mei, Xin Yang 0011, Yang Wang 0106, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau
CVPR2
2020 Attention-Guided Hierarchical Structure Aggregation for Image Matting
abstract
Existing deep learning based matting algorithms primarily resort to high-level semantic features to improve the overall structure of alpha mattes. However, we argue that advanced semantics extracted from CNNs contribute unequally for alpha perception and we are supposed to reconcile advanced semantic information with low-level appearance cues to refine the foreground details. In this paper, we propose an end-to-end Hierarchical Attention Matting Network (HAttMatting), which can predict the better structure of alpha mattes from single RGB images without additional input. Specifically, we employ spatial and channel-wise attention to integrate appearance cues and pyramidal features in a novel fashion. This blended attention mechanism can perceive alpha mattes from refined boundaries and adaptive semantics. We also introduce a hybrid loss function fusing Structural SIMilarity (SSIM), Mean Square Error (MSE) and Adversarial loss to guide the network to further improve the overall foreground structure. Besides, we construct a large-scale image matting dataset comprised of 59,600 training images and 1000 test images (total 646 distinct foreground alpha mattes), which can further improve the robustness of our hierarchical structure aggregation model. Extensive experiments demonstrate that the proposed HAttMatting can capture sophisticated foreground structure and achieve state-of-the-art performance with single RGB images as input.
Yu Qiao 0001, Yuhao Liu 0001, Xin Yang 0011, Mingliang Xu 0001, Qiang Zhang 0008, Xiaopeng Wei
CVPR3
2020 TENet: Triple Excitation Network for Video Salient Object Detection
Sucheng Ren, Chu Han, Xin Yang 0011, Guoqiang Han 0002, Shengfeng He
ECCV (5)3
2020 Kernel Clustering On Symmetric Positive Definite Manifolds Via Double Approximated Low Rank Representation
abstract
As an effective descriptor, Symmetric Positive Definite (SPD) matrix is widely used in several areas such as image clustering. Recently, researchers proposed some effective methods based on low rank theory for SPD data clustering with nonlinear metric. However, single nuclear norm is always adopted to formulate the low rank model in these methods, which would lead to suboptimal solution. In this paper, we proposed a novel double low rank representation method for SPD clustering problem, in which matrix factorization and nonconvex rank constraint are combined to reveal the intrinsic property of the data instead of employing the nuclear norm. Meanwhile, kernel method and Log-Euclidean metric are combined to better explore the intrinsic geometry within SPD data. The proposed method has been evaluated on several public datasets and the experimental results demonstrate that the proposed method outperforms the state-of-the-art ones.
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011, Wenwu Zhu 0001, Ge Li 0002
ICME5
2020 Multi-Context And Enhanced Reconstruction Network For Single Image Super Resolution
abstract
Most existing single image super-resolution (SISR) methods continually increase the depth or width of networks, without adequately exploring contextual features which are essential for reconstruction. Moreover, such existing methods pay little attention to the final high-resolution(HR) image reconstruction step and therefore hinder the desired SR performance. In this paper, we propose a multi-context and enhanced reconstruction network (MCERN) for SISR. Specifically, a novel model named Multi-Context Block (MCB) which extracts more image contextual features with multibranch dilated convolution. Applying multiple MCBs with residual and dense connections, we can effectively extract contextual and hierarchical features for obtaining the coarse super-resolution result. Then an enhanced reconstruction block (ERB) is followed to extract essential spatial features on the high-resolution image to refine the coarse result to a better result. Extensive benchmark evaluations demonstrate the efficacy of our proposed MCERN in terms of metric accuracy and visual effects.
Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Xin Yang 0011, Haiyang Mei
ICME4
2020 A Spectral Clustering on Grassmann Manifold via Double Low Rank Constraint
abstract
Data clustering is a fundamental topic in machine learning and data mining areas. In recent years, researchers have proposed a series of effective methods based on Low Rank Representation (LRR) which could explore low-dimension subspace structure embedded in original data effectively. The traditional LRR methods usually are designed for vectorial data from linear spaces with Euclidean distance. However, high-dimension data (such as video clip or imageset) are always considered as non-linear manifold data such as Grassmann manifold with non-linear metric. In addition, traditional LRR clustering method always adopt single nuclear norm as low rank constraint which would lead to suboptimal solution and decrease the clustering accuracy. In this paper, we proposed a new low rank method on Grassmann manifold for video or imageset data clustering task. In the proposed method, video or imageset data are formulated as sample data on Grassmann manifold first. And then a double low rank constraint is proposed by combining the nuclear norm and bilinear representation for better construct the representation matrix. The experimental results on several public datasets show that the proposed method outperforms the state-of-the-art clustering methods.
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011
ICPR5
2020 A Preliminary Study of Fusion ARTs with Adaptively Information Intensity Attenuation Controlling
abstract
Fusion ART is an enhanced version of Adaptive Resonance Theory (ART) which is derived from a biologically-plausible theory of human cognitive information processing. Due to its well-established ability of learning associative mappings across multimodal pattern channels in an online and incremental manner, fusion ART has been widely applied in many real world learning problems. In this paper, we take a Fusion Architecture for Learning, Cognition, and Navigation (FALCON) as the specification and essential backbone of fusion ART and introduce an intensity attenuation controller δ for adaptively adjusting the intensity of information captured from the environment, by taking inspiration from Broadbent-Treisman Filter-Attenuation's perceptual model of environmental attention. Particularly, we propose both an adaptive δ detection algorithm as well as a δ-based pruning algorithm to enhance the learning performance of FALCON while reduce the redundant memory storage incurred by the "detrimental δ". To verify the effectiveness and efficiency of our proposed method, comprehensive experimental studies are carried out on a classical minefield navigation task.
Wenxuan Zhu, Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Liang Feng 0001, Xinghua Qu
IJCNN5
2020 Multi-scale Information Assembly for Image Matting
abstract
Abstract Image matting is a long‐standing problem in computer graphics and vision, mostly identified as the accurate estimation of the foreground in input images. We argue that the foreground objects can be represented by different‐level information, including the central bodies, large‐grained boundaries, refined details, etc. Based on this observation, in this paper, we propose a multi‐scale information assembly framework (MSIA‐matte) to pull out high‐quality alpha mattes from single RGB images. Technically speaking, given an input image, we extract advanced semantics as our subject content and retain initial CNN features to encode different‐level foreground expression, then combine them by our well‐designed information assembly strategy. Extensive experiments can prove the effectiveness of the proposed MSIA‐matte, and we can achieve state‐of‐the‐art performance compared to most existing matting networks.
Yu Qiao 0001, Yuhao Liu 0001, Xin Yang 0011, Yuxin Wang 0001, Qiang Zhang 0008, Xiaopeng Wei
Comput. Graph. Forum4
2020 Point cloud semantic scene segmentation based on coordinate convolution
abstract
Abstract Point cloud semantic segmentation, a crucial research area in the 3D computer vision, lies at the core of many vision and robotics applications. Due to the irregular and disordered of the point cloud, however, the application of convolution on point clouds is challenging. In this article, we propose the “coordinate convolution,” which can effectively extract local structural information of the point cloud, to solve the inapplicability of conventional convolution neural network (CNN) structures on the 3D point cloud. The “coordinate convolution” is a projection operation of three planes based on the local coordinate system of each point. Specifically, we project the point cloud on three planes in the local coordinate system with a joint 2D convolution operation to extract its features. Additionally, we leverage a self‐encoding network based on image semantic segmentation U‐Net structure as the overall architecture of the point cloud semantic segmentation algorithm. The results demonstrate that the proposed method exhibited excellent performances for point cloud data sets corresponding to various scenes.
Zhaoxuan Zhang, Xuefeng Yin, Xinglin Piao, Yuxin Wang 0001, Xin Yang 0011
Comput. Animat. Virtual Worlds6
2020 RGB-D salient object detection via deep fusion of semantics and details
abstract
Abstract In this paper, we address RGB‐D salient object detection task by jointly leveraging semantics and contour details of salient objects. We propose a novel semantics‐and‐details complementary fusion network to adaptively integrate cross‐model and multilevel features. Specifically, we employ two kinds of fusion modules in our model, which are designed for fusing high‐level semantic features and integrating contour detail features of the scene components, respectively. The semantics fusion module aggregates high‐level interdependent semantic relationships by a nonlinear weighted summation of small and medium receptive fields. Meanwhile, the details module integrates multi‐level contour detail features to leverage expressive details of salient objects. We achieve new state‐of‐the‐art salient object detection results on seven RGB‐D datasets, that is, STERE, NJU2000, LFSD, NLPR, SSD, DES, and SIP2019 dataset. Experimental results demonstrate that our method outperforms eleven state‐of‐the‐art salient object detection methods.
Shimin Zhao, Pengjie Wang 0001, Ying Cao 0001, Xin Yang 0011
Comput. Animat. Virtual Worlds6
2019 Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image
abstract
We present a deep reinforcement learning method of progressive view inpainting for 3D point scene completion under volume guidance, achieving high-quality scene reconstruction from only a single depth image with severe occlusion. Our approach is end-to-end, consisting of three modules: 3D scene volume reconstruction, 2D depth map inpainting, and multi-view selection for completion. Given a single depth image, our method first goes through the 3D volume branch to obtain a volumetric scene reconstruction as a guide to the next view inpainting step, which attempts to make up the missing information; the third step involves projecting the volume under the same view of the input, concatenating them to complete the current view depth, and integrating all depth into the point cloud. Since the occluded areas are unavailable, we resort to a deep Q-Network to glance around and pick the next best view for large hole completion progressively until a scene is adequately reconstructed while guaranteeing validity. All steps are learned jointly to achieve robust and consistent results. We perform qualitative and quantitative evaluations with extensive experiments on the SUNCG data, obtaining better results than the state of the art.
Xiaoguang Han 0001, Zhaoxuan Zhang, Dong Du 0002, Mingdai Yang, Jingming Yu, Xin Yang 0011, Ligang Liu 0001, Zixiang Xiong, Shuguang Cui
CVPR7
2019 Spatial Attentive Single-Image Deraining With a High Quality Real Rain Dataset
abstract
Removing rain streaks from a single image has been drawing considerable attention as rain streaks can severely degrade the image quality and affect the performance of existing outdoor vision tasks. While recent CNN-based derainers have reported promising performances, deraining remains an open problem for two reasons. First, existing synthesized rain datasets have only limited realism, in terms of modeling real rain characteristics such as rain shape, direction and intensity. Second, there are no public benchmarks for quantitative comparisons on real rain images, which makes the current evaluation less objective. The core challenge is that real world rain/clean image pairs cannot be captured at the same time. In this paper, we address the single image rain removal problem in two ways. First, we propose a semi-automatic method that incorporates temporal priors and human supervision to generate a high-quality clean image from each input sequence of real rain images. Using this method, we construct a large-scale dataset of ∼29.5K rain/rain-free image pairs that covers a wide range of natural rain scenes. Second, to better cover the stochastic distribution of real rain streaks, we propose a novel SPatial Attentive Network (SPANet) to remove rain streaks in a local-to-global manner. Extensive experiments demonstrate that our network performs favorably against the state-of-the-art deraining methods.
Tianyu Wang 0003, Xin Yang 0011, Ke Xu 0010, Shaozhe Chen, Qiang Zhang 0008, Rynson W. H. Lau
CVPR2
2019 Where Is My Mirror?
Xin Yang 0011, Haiyang Mei, Ke Xu 0010, Xiaopeng Wei, Rynson W. H. Lau
ICCV1
2019 Real-virtual consistent traffic flow interaction
Xin Yang 0011, Shuai Li 0014, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei
Graph. Model.1
2019 DEMC: A Deep Dual-Encoder Network for Denoising Monte Carlo Rendering
Xin Yang 0011, Wenbo Hu 0002, Lijing Zhao, Qiang Zhang 0008, Xiaopeng Wei, Hongbo Fu 0001
J. Comput. Sci. Technol.1
2019 Cascaded network with deep intensity manipulation for scene understanding
abstract
Abstract Scene understanding is essential to robotic navigation and autonomous driving as it provides semantic information to their controlling system. However, it will fail when processing low‐light images/videos captured under adverse weather or at night use state‐of‐the‐art scene understanding methods. A naive way to directly infer semantics from low‐light images is ill posed because the low‐light condition distorts pixel intensities and buries details. In order to address this problem, we propose the Deep Intensity Manipulation Network (DIMNet), which could relight the input images and recover the details, and combine the DIMNet with a scene understanding network to get a cascaded network to learn the semantics from low‐light images. Through learning pixel intensity manipulation, our method can generate images not only visually pleasing but also practical for scene understanding. Qualitative and quantitative experiments demonstrate that the proposed method is effective and robust for both synthetic and real‐world images.
Xin Yang 0011, Shaozhe Chen, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei
Comput. Animat. Virtual Worlds1
2019 DRFN: Deep Recurrent Fusion Network for Single-Image Super-Resolution With Large Factors
abstract
Recently, single-image super-resolution has made great progress due to the development of deep convolutional neural networks (CNNs). The vast majority of CNN-based models use a predefined upsampling operator, such as bicubic interpolation, to upscale input low-resolution images to the desired size and learn nonlinear mapping between the interpolated image and ground truth high-resolution (HR) image. However, interpolation processing can lead to visual artifacts as details are over smoothed, particularly when the super-resolution factor is high. In this paper, we propose a deep recurrent fusion network (DRFN), which utilizes transposed convolution instead of bicubic interpolation for upsampling and integrates different-level features extracted from recurrent residual blocks to reconstruct the final HR images. We adopt a deep recurrence learning strategy and, thus, have a larger receptive field, which is conducive to reconstructing an image more accurately. Furthermore, we show that the multilevel fusion structure is suitable for dealing with image super-resolution problems. Extensive benchmark evaluations demonstrate that the proposed DRFN performs better than most current deep learning methods in terms of accuracy and visual effects, especially for large-scale images, while using fewer parameters.
Xin Yang 0011, Haiyang Mei, Jiqing Zhang, Ke Xu 0010, Qiang Zhang 0008, Xiaopeng Wei
IEEE Trans. Multim.1
2018 Image Correction via Deep Reciprocating HDR Transformation
abstract
Image correction aims to adjust an input image into a visually pleasing one. Existing approaches are proposed mainly from the perspective of image pixel manipulation. They are not effective to recover the details in the under/over exposed regions. In this paper, we revisit the image formation procedure and notice that the missing details in these regions exist in the corresponding high dynamic range (HDR) data. These details are well perceived by the human eyes but diminished in the low dynamic range (LDR) domain because of the tone mapping process. Therefore, we formulate the image correction task as an HDR transformation process and propose a novel approach called Deep Reciprocating HDR Transformation (DRHT). Given an input LDR image, we first reconstruct the missing details in the HDR domain. We then perform tone mapping on the predicted HDR data to generate the output LDR image with the recovered details. To this end, we propose a united framework consisting of two CNNs for HDR reconstruction and tone mapping. They are integrated end-to-end for joint training and prediction. Experiments on the standard benchmarks demonstrate that the proposed method performs favorably against state-of-the-art image correction methods.
Xin Yang 0011, Ke Xu 0010, Yibing Song, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau
CVPR1
2018 Active Object Reconstruction Using a Guided View Planner
abstract
Inspired by the recent advance of image-based object reconstruction using deep learning, we present an active reconstruction model using a guided view planner. We aim to reconstruct a 3D model using images observed from a planned sequence of informative and discriminative views. But where are such informative and discriminative views around an object? To address this we propose a unified model for view planning and object reconstruction, which is utilized to learn a guided information acquisition model and to aggregate information from a sequence of images for reconstruction. Experiments show that our model (1) increases our reconstruction accuracy with an increasing number of views (2) and generally predicts a more informative sequence of views for object reconstruction compared to other alternative methods.
Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei, Hongbo Fu 0001
IJCAI1
2018 Active Matting
abstract
Image matting is an ill-posed problem. It requires a user input trimap or some strokes to obtain an alpha matte of the foreground object. A fine user input is essential to obtain a good result, which is either time consuming or suitable for experienced users who know where to place the strokes. In this paper, we explore the intrinsic relationship between the user input and the matting algorithm to address the problem of where and when the user should provide the input. Our aim is to discover the most informative sequence of regions for user input in order to produce a good alpha matte with minimum labeling efforts. To this end, we propose an active matting method with recurrent reinforcement learning. The proposed framework involves human in the loop by sequentially detecting informative regions for trivial human judgement. Comparing to traditional matting algorithms, the proposed framework requires much less efforts, and can produce satisfactory results with just 10 regions. Through extensive experiments, we show that the proposed model reduces user efforts significantly and achieves comparable performance to dense trimaps in a user-friendly manner. We further show that the learned informative knowledge can be generalized across different matting algorithms.
Xin Yang 0011, Ke Xu 0010, Shaozhe Chen, Shengfeng He, Rynson W. H. Lau
NeurIPS1
2018 Efficient image super-resolution integration
Ke Xu 0010, Xin Wang 0118, Xin Yang 0011, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau
Vis. Comput.3
2017 Real-virtual fusion model for traffic animation
abstract
Abstract In this paper, we present an innovative, animated traffic simulation method that we designed to feature an enhanced sense of reality and diversity of traffic flows. Instead of the typical one‐off initialization, our simulation method includes continuous, real trajectory data input providing an interactive control function that maximizes the characteristics of real‐world traffic flows. Our fusion models represent a comprehensive integration of the interactions among real‐data‐driven and virtual vehicles, thus depicting accurately the irregularity of traffic flows. Test results showed that animations generated via our proposed method depict inverse and irregular vehicle driving behaviors throughout the entire traffic flow.
Xin Yang 0011, Wanchao Su, Xiaogang Jin 0001, Guozhen Tan
Comput. Animat. Virtual Worlds1
2017 Interactive traffic simulation model with learned local parameters
Xin Yang 0011, Shuai Li 0014, Wanchao Su, Guozhen Tan, Qiang Zhang 0008, Xiaopeng Wei
Multim. Tools Appl.1
2017 3D palmprint recognition using shape index representation and fragile bits
Xueqin Xiang, Duanqing Xu, Xin Yang 0011
Multim. Tools Appl.5
2016 DKD: a fast k-d tree update design for dynamic scenes
abstract
Abstract We design dynamic k‐d (DKD) tree based on classical k‐d tree for animated scene rendering. Our method can inherit the benefit of efficient traversal of k‐d tree and minimize time cost to update DKD tree, making it well suited for animated geometry. DKD employs primitive reset, redistribution to reflect the updated positions of geometry, and leaf node incremental growing to avoid the deterioration of hierarchy quality due to refitting. Our experiments show that DKD has a significant rendering performance improvement than selected existing methods. Copyright © 2016 John Wiley & Sons, Ltd.
Xin Yang 0011, Pengfei Zhang 0016, Lutong Xin, Yuxin Wang 0001, Qiang Zhang 0008, Xiaopeng Wei
Comput. Animat. Virtual Worlds1
2016 MSKD: multi-split KD-tree design on GPU
Xin Yang 0011, Pengjie Wang 0001, Duanqing Xu
Multim. Tools Appl.1
2015 Complex shading efficiently for ray tracing on GPU
Xin Yang 0011, Duanqing Xu
Multim. Tools Appl.1
2014 Optimized big data K-means clustering using MapReduce
Xiaoli Cui, Pingfei Zhu, Xin Yang 0011, Keqiu Li, Changqing Ji
J. Supercomput.3