Wei Zhang 0021

dblp:10/4661-21 · DBLP profile ↗
← Back
146ranked-venue papers
28as first author
75since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 17 first-author · 29 since 2021Artificial intelligence and machine learning · 67 · 10 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 21 since 2021Systems, architecture and hardware · 13 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A whole-life fatigue crack growth rate prediction method based on active learning and physics-informed loss
Qixuan Zhang, Wei Zhang 0021, Rui Huang 0001, Xinghui Chen, Changyu Zhou
Eng. Appl. Artif. Intell.2
2026 Breaking Barriers, Localizing Saliency: A Large-Scale Benchmark and Baseline for Condition-Constrained Salient Object Detection
abstract
Salient Object Detection (SOD) aims to identify and segment the most prominent objects in an image. In real open environments, intelligent systems often encounter complex and challenging scenes, such as low-light, rain, snow, etc., which we call constrained conditions. These real situations pose more severe challenges to existing SOD models. However, there is no comprehensive and in-depth exploration of this field at both the data and model levels, and most of them focus on ideal situations or a single condition. To bridge this gap, we launch a new task, Condition-Constrained Salient Object Detection (CSOD), aimed at robustly and accurately locating salient objects in constrained environments. On the one hand, to compensate for the lack of datasets, we construct the first large-scale condition-constrained salient object detection dataset CSOD10 K, comprising 10,000 pixel-level annotated images and over 100 categories of salient objects. This dataset is oriented towards the real environment and includes 8 real-world constrained scenes under 3 main constraint types, making it extremely challenging. On the other hand, we abandon the paradigm of "restoration before detection" and instead introduce a unified end-to-end framework CSSAM that fully explores scene attributes, eliminating the need for additional ground-truth restored images and reducing computational overhead. Specifically, we design a Scene Prior-Guided Adapter (SPGA), which injects scene priors to enable the foundation model to better adapt to downstream constrained scenes. To automatically decode salient objects, we propose a Hybrid Prompt Decoding Strategy (HPDS), which can effectively integrate multiple types of prompts to achieve adaptation to the SOD task. Extensive experiments show that our model significantly outperforms state-of-the-art methods on both the CSOD10 K dataset and existing standard SOD benchmarks.
Runmin Cong, Hao Fang 0010, Sam Kwong, Wei Zhang 0021
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Dexterous Manipulation Through Imitation Learning: A Survey
abstract
Dexterous manipulation, which refers to the ability of a robotic hand or multi-fingered end-effector to skillfully control, reorient, and manipulate objects through precise, coordinated finger movements and adaptive force modulation, enables complex interactions similar to human hand dexterity. With recent advances in robotics and machine learning, there is a growing demand for these systems to operate in complex and unstructured environments. Traditional model-based approaches struggle to generalize across tasks and object variations due to the high dimensionality and complex contact dynamics of dexterous manipulation. Although model-free methods such as reinforcement learning (RL) show promise, they require extensive training, large-scale interaction data, and carefully designed rewards for stability and effectiveness. Imitation learning (IL) offers an alternative by allowing robots to acquire dexterous manipulation skills directly from expert demonstrations, capturing fine-grained coordination and contact dynamics while bypassing the need for explicit modeling and large-scale trial-and-error. This survey provides an overview of dexterous manipulation methods based on imitation learning, details recent advances, and addresses key challenges in the field. Additionally, it explores potential research directions to enhance IL-driven dexterous manipulation. Our goal is to offer researchers and practitioners a comprehensive introduction to this rapidly evolving domain.
Shan An, Chao Tang 0001, Yuning Zhou, Tengyu Liu, Fangqiang Ding, Shufang Zhang, Yao Mu 0001, Ran Song 0001, Wei Zhang 0021, Zeng-Guang Hou, Hong Zhang 0013
IEEE Trans Autom. Sci. Eng.10
2026 MoTL: Modality-Balanced Terrain-Aware Locomotion Learning for Quadruped Robots
Zhiheng Li 0005, Yanyun Chen, Wenhao Tan, Mingxin Zhang 0006, Ran Song 0001, Wei Zhang 0021
IEEE Trans Autom. Sci. Eng.7
2026 PAPNet: Point-Enhanced Attention-Aware Pillar Network for 3D Object Detection in Autonomous Driving
abstract
The conversion of raw point clouds into pillar representations has been widely adopted for 3D object detection. Such conversion allows a point cloud to be discretized into structured grids, which enables more efficient spatial representation and faster processing in real-time autonomous driving systems. However, discretizing raw point clouds often leads to the misdetection of small objects such as pedestrians and cyclists. This is because the discretization inevitably results in the loss of contextual and multi-resolution information within raw point clouds. To address this issue, we propose PAPNet, a point-enhanced attention-aware pillar network mainly composed of a point-pillar cross-attention module (PCM), a pillar-wise dual attention module (PDAM), and a multi-resolution set abstraction module (MSAM). PCM integrates raw point cloud features with pillar features across different dimensions, and PDAM guides PAPNet to focus on the intrinsic characteristics of the pillars. Additionally, MSAM retains both high-resolution and low-resolution features while integrating multi-scale information. Extensive experiments on four public datasets and in real-world scenarios demonstrate the effectiveness and efficiency of PAPNet. Codes, data, and demo videos can be found at the project website https://vsislab.github.io/PAPNet/.
Ruitong Li, Yuenan Zhao, Jiaming Chen 0001, Ran Song 0001, Wei Zhang 0021
IEEE Trans Autom. Sci. Eng.6
2026 CoVA-IL: Zero-Shot Imitation Learning via Contrastive Viewpoint Alignment on Object-Centric Representation
abstract
Imitation learning provides an efficient paradigm for acquiring robotic manipulation skills, yet policies trained in a single environment often generalize poorly to unseen scenes. To address this challenge, we propose CoVA-IL, a zero-shot imitation learning framework that performs contrastive viewpoint alignment on object-centric representations, enabling direct policy deployment in novel environments without retraining. CoVA-IL uses the target object’s point cloud as the visual input and learns viewpoint-invariant latent representations through contrastive learning, thereby improving robustness to background changes, viewpoint variations, and cross-environment shifts. In addition, we incorporate multi-level point-cloud augmentation into 3D visuomotor imitation learning to improve data efficiency and reduce the number of demonstrations required for training. Real-world experiments show that CoVA-IL maintains average task success rates of 82.5% and 80% under substantial background and viewpoint changes, respectively, and multi-scene evaluations further validate its effectiveness for cross-environment deployment. Moreover, CoVA-IL can learn basic manipulation skills from a small number of real demonstrations, thereby reducing demonstration collection costs.
Jiangtao Luo, Chenchen Zheng, Jinqiu Fan, Ran Song 0001, Wei Zhang 0021
IEEE Trans Autom. Sci. Eng.7
2026 RDPrompter: Reference-Defect Prompt Learning for Few-Shot Defect Segmentation Based on Visual Foundation Model
abstract
Few-shot industrial defect segmentation (FIDS) is an extremely challenging task in industrial inspection, which focuses on segmenting unseen defect categories with only a few samples. Existing few-shot segmentation methods are commonly developed with constrained feature extraction capabilities, making it difficult to accurately segment diverse unseen defects from distinctive industrial scenarios. To this end, we aim to leverage strong generalization strengths of the large-scale visual foundation model, segment anything model (SAM), to handle FIDS task. However, simply incorporating SAM into FIDS is insufficient for automatically categorizing defects, as the SAM heavily relies on manual prompts for segmentation. To address this, we propose a novel reference-defect prompt learning method RDPrompter, which learns adaptive prompts from defect similarities based on SAM for guiding automated FIDS. Specifically, to obtain adaptive prompts, we introduce a multiscale feature extraction strategy to mine rich defect features by the image encoder of SAM. Then, self-correcting probability prototype and cosine similarity attention are proposed to form a mixed similarity aggregation module, which obtains multiscale feature similarities for defect localization. Based on the similarities, a prompt embedding construction module is designed to further extract fine defect location information and generate prompt embeddings for FIDS. Sufficient experiments on MVTec-Unseen, SDD, FSSD-12 and the collected CID datasets demonstrate the superiority of our method.
Tiyu Fang, Lin Zhang 0041, Ran Song 0001, Xiaolei Li 0003, Wei Zhang 0021
IEEE Trans. Ind. Informatics6
2026 DiffLLFace: Learning Alternate Illumination-Diffusion Adaptation for Low-Light Face Super-Resolution and Beyond
abstract
Facial image acquisition under constrained illumination and with limited-resolution imaging devices often results in coupled photometric and geometric degradations, manifesting as low-light and low-resolution (LLR) conditions. Prevailing research predominantly follows fragmented optimization paradigms that address low-light image enhancement (LLIE) and face super-resolution (FSR) as isolated tasks. This approach overlooks the compound nature of the degradations, thereby significantly limiting their applicability in practical scenarios. To bridge this gap, we present DiffLLFace, a unified framework that harnesses diffusive generative capabilities with illumination-aware trajectories to achieve robust FSR from LLR observations. The core of our method lies in its alternate illumination-diffusion adaptation, which operates throughout the generation process. This mechanism not only captures degradation patterns in both brightness and structure to harmonize latent representations but also dynamically calibrates the illumination prior with the generative knowledge inherent to diffusion models. As such, DiffLLFace attains precise control over conditional adaptation and illumination rectification. We further devise a simple yet effective non-parametric Fourier enhancement strategy, which provides structural appearance clues that work in concert with the alternate adaptation to ensure texture and color consistency. Extensive experiments demonstrate the superiority of DiffLLFace over existing methods and remarkable generalizability on complex natural scenes. Code is available at https://github.com/KaishengPang/DiffLLFace.
Runmin Cong, Kaisheng Pang, Feng Li 0037, Hua Li 0012, Huihui Bai 0001, Sam Kwong, Wei Zhang 0021
IEEE Trans. Image Process.7
2026 CVDII: Enhancing One-Shot Skeleton Action Recognition Through Cross-View Dynamic Information Interaction
abstract
One-shot 3D skeleton action recognition task struggles with diverse intra-class action execution styles, causing excessive discriminative information to obstruct obtaining separable feature space. We innovatively propose leveraging shared information among intra-class action executions to mitigate the over-influence of discriminative information. To this end, we proposed dynamic information interaction module (DIIM) that enables shared information to effectively weaken excessive discriminative information. Specifically, DIIM facilitates effective information interaction by constructing a guided evolution pool to store execution-related shared information and ensure such information can be retrieved. We devise shared-discriminative projection strategy (SDPS) which adopts different feature extraction strategies for specific skeleton topologies to target mining discriminative and shared information from different views of skeleton data. In summary, our proposed Cross-View Dynamic Information Interaction (CVDII) framework integrates DIIM and SDPS, effectively tackles the problem of discriminative information redundancy caused by diverse intra-class action execution styles. Experiments conducted on NTU 60, NTU 120, PKU-MMD, and Kinetics datasets demonstrate that our proposed CVDII achieves remarkable performance.
Youmei Zhang, Weidong Zhang 0005, Zhiheng Li 0005, Bin Li 0042, Wei Zhang 0021
IEEE Trans. Image Process.7
2026 B2Q-Net: Bidirectional Branch Query Network for Surgical Phase Recognition
abstract
Surgical phase recognition (SPR) is essential for surgical workflow analysis and provides immediate guidance during procedures. Existing methods aggregate frame-level information into a global representation and treat the task as frame-wise classification. However, this pipeline lacks a feedback mechanism for integrating historical information into local temporal modeling. To address this limitation, we propose the Bidirectional Branch Query Network (B2Q-Net), which reformulates the SPR task as the bidirectional query between phase-level features and frame-level features. B2Q-Net incorporates historical information during the initialization of phase queries. This enables bidirectional information flow during iterative refinement of two-level feature maps between phases and frames. Furthermore, we introduce a dual-scale selector (DSS) to generate high-quality phase queries for the current video clip. These phase queries retrieve historical information from the proposed state space query (SSQ) module, which uses learnable tokens as the historical state space to preserve historical information. Extensive evaluations on three datasets demonstrate that B2Q-Net consistently outperforms state-of-the-art methods in recognition accuracy while achieving an inference speed of 106 fps. The B2Q-Net code is available at https://github.com/vsislab/B2Q-Net.
Zhiheng Li 0005, Yue Bi, Xiao Jia 0005, Ran Song 0001, Wei Zhang 0021
IEEE Trans. Medical Imaging7
2025 Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction Tuning
abstract
Recent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-ranging code-related tasks. However, most previous existing methods mainly view each programming language in isolation and ignore the knowledge transfer among different programming languages. To bridge the gap among different programming languages, we introduce a novel multi-agent collaboration framework to enhance multilingual instruction tuning for code LLMs, where multiple language-specific intelligent agent components with generation memory work together to transfer knowledge from one language to another efficiently and effectively. Specifically, we first generate the language-specific instruction data from the code snippets and then provide the generated data as the seed data for language-specific agents. Multiple language-specific agents discuss and collaborate to formulate a new instruction and its corresponding solution (A new programming language or existing programming language), To further encourage the cross-lingual transfer, each agent stores its generation history as memory and then summarizes its merits and faults. Finally, the high-quality multilingual instruction data is used to encourage knowledge transfer among different programming languages to train Qwen2.5-xCoder. Experimental results on multilingual programming benchmarks demonstrate the superior performance of Qwen2.5-xCoder in sharing common knowledge, highlighting its potential to reduce the cross-lingual gap.
Jian Yang 0003, Wei Zhang 0021, Yibo Miao, Shanghaoran Quan, Zhenhe Wu, Qiyao Peng 0006, Liqun Yang, Tianyu Liu 0001, Zeyu Cui, Binyuan Hui, Junyang Lin
ACL (1)2
2025 Decoupled Motion Expression Video Segmentation
abstract
Motion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is more challenging. Previous works achieved it by simply injecting text information into the video instance segmentation (VIS) model. However, this requires retraining the entire model and optimization is difficult. In this work, we propose DMVS, a simple framework constructed on the existing query-based VIS model, emphasizing decoupling the task into video instance segmentation and motion expression understanding. Firstly, we use a frozen video instance segmenter to extract object-specific contexts and convert them into frame-level and video-level queries. Secondly, we interact two levels of queries with static and motion cues, respectively, to further encode visually enhanced motion expressions. Furthermore, we propose a novel query initialization strategy that uses video queries guided by classification priors to initialize motion queries, greatly reducing the difficulty of optimization. Without bells and whistles, DMVS achieves state-of-the-art performance on the MeViS dataset at a lower training cost. Extensive experiments verify the effectiveness and efficiency of our framework.
Hao Fang 0010, Runmin Cong, Xiankai Lu, Xiaofei Zhou 0003, Sam Kwong, Wei Zhang 0021
CVPR6
2025 GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill
abstract
Learning open-vocabulary physical skills for simulated agents presents a significant challenge in Artificial Intelligence (AI). Current Reinforcement Learning (RL) approaches face critical limitations: manually designed rewards lack scalability across diverse tasks, while demonstration-based methods struggle to generalize beyond their training distribution. We introduce GROVE, a generalized reward framework that enables open-vocabulary physical skill learning without manual engineering or task-specific demonstrations. Our key insight is that Large Language Models (LLMs) and Vision Language Models (VLMs) provide complementary guidance—LLMs generate precise physical constraints capturing task requirements, while VLMs evaluate motion semantics and naturalness. Through an iterative design process, VLM-based feedback continuously refines LLM-generated constraints, creating a self-improving reward system. To bridge the domain gap between simulation and natural images, we develop Pose2CLIP, a lightweight mapper that efficiently projects agent poses directly into semantic feature space without computationally expensive rendering. Extensive experiments across diverse embodiments and learning paradigms demonstrate GROVE’s effectiveness, achieving 22.2% higher motion naturalness and 25.7% better task completion scores while training 8.4× faster than previous methods. These results establish a new foundation for scalable physical skill acquisition in simulated environments.
Jieming Cui, Tengyu Liu, Jiale Yu, Ran Song 0001, Wei Zhang 0021, Yixin Zhu 0001, Siyuan Huang 0001
CVPR6
2025 CodeArena: Evaluating and Aligning CodeLLMs on Human Preference
abstract
Jian Yang, Jiaxi Yang, Wei Zhang, Jin Ke, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li, Binyuan Hui, Junyang Lin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jian Yang 0003, Jiaxi Yang 0004, Wei Zhang 0021, Yibo Miao, Lei Zhang 0201, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li 0001, Binyuan Hui, Junyang Lin
EMNLP3
2025 ALVO: Adaptive Learning with Velocity Obstacles for UGV Navigation in Dynamic Scenes
abstract
Autonomous navigation of unmanned ground vehicles (UGVs) in dynamic scenes is a challenging task that requires them to avoid obstacles and move toward the goal simultaneously. This paper proposes ALVO, an adaptive learning policy that leverages velocity obstacles for UGV navigation. ALVO employs an adaptive gating-based mechanism for reactive obstacle avoidance, which enables the UGV to either slow down or proactively navigate around obstacles based on the relative importance of the environmental state and the goal. A reward function based on velocity obstacles is also designed to guide the UGV to navigate toward the goal while avoiding obstacles. Extensive experiments demonstrate that ALVO outperforms the competing approaches in various dynamic environments. We also implemented our method on a real UGV and showed that it performed well in real-world scenarios.
Yinduo Xie, Yuenan Zhao, Ran Song 0001, Zhiheng Li 0005, Lei Han 0001, Wei Zhang 0021
IROS6
2025 VCADNet: Vision-based Circular Accessible Depth Prediction for UGV Perception
abstract
Circular accessible depth (CAD) provides a lightweight and robust traversability representation for autonomous navigation of unmanned ground vehicles (UGV). Aiming at the limitations of existing LiDAR-based methods in detecting low-thickness targets and executing semantic reasoning, we propose VCADNet, a vision-based neural network for circular accessible depth prediction. VCADNet comprises three core components: a geometry-based query module for multi-view bird’s eye view feature extraction, a polar coordinate transformation for CAD alignment, and a multi-scale U-Net architecture for depth prediction. In addition, we present a cross-modal contrastive learning scheme to enhance the spatial reasoning of VCADNet, which transfers knowledge from LiDAR-based encoders to vision-based counterparts. Extensive experiments demonstrate the superior performance of VCADNet in various UGV perception tasks.
Yuenan Zhao, Ran Song 0001, Lei Han 0001, Wei Zhang 0021
IROS6
2025 KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
abstract
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.
Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009
NeurIPS15
2025 Learning packing-and-unpacking synergistic policy via LLM-guided DRL for robust online robotic packing
Shuai Song, Ran Song 0001, Jiyu Cheng, Yibin Li 0001, Wei Zhang 0021
Adv. Eng. Informatics6
2025 C2P-Net: Comprehensive Depth Map to Planar Depth Conversion for Room Layout Estimation
abstract
Room layout estimation seeks to infer the overall spatial configuration of indoor scenes using perspective or panoramic images. As the layout is determined by the dominant indoor planes, this problem inherently requires the reconstruction of these planes. Some studies reconstruct indoor planes from perspective images by learning pixel-level or instance-level plane parameters. However, directly learning these parameters has the problems of susceptibility to occlusions and position dependency. In this paper, we introduce the Comprehensive depth map to Planar depth (C2P) conversion, which reformulates planar depth reconstruction into the prediction of a comprehensive depth map and planar visibility confidence. Based on the parametric representation of planar depth we propose, the C2P conversion is applicable to both panoramic and perspective images. Accordingly, we present an effective framework for room layout estimation that jointly learns the comprehensive depth map and planar visibility confidence. Due to the differentiability of the C2P conversion, our network autonomously learns planar visibility confidence by constraining the estimated plane parameters and reconstructed planar depth map. We further propose a novel approach for 3D layout generation through sequential planar depth map integration. Experimental results demonstrate the superiority of our method across all evaluated panoramic and perspective datasets.
Weidong Zhang 0005, Mengjie Zhou, Jiyu Cheng, Ying Liu 0026, Wei Zhang 0021
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 A Generalist Agent Learning Architecture for Versatile Quadruped Locomotion
abstract
Quadrupeds can generate various motor behaviors with the muscle synergies activated by the central nervous system. However, versatile locomotion for quadruped robots remains challenging due to the complexity of the high-dimensional limb dynamics with many physical constraints. Current approaches typically apply a dedicated policy or controller for each motor behavior, which requires the optimization of a large number of parameters and training process is complicated. In this paper, we propose a Generalist Agent Learning Architecture (GALA) to learn diverse motor behaviors simultaneously with a single policy network for quadruped locomotion. GALA significantly decreases the number of trainable parameters while producing appropriate motor behaviors by simply reactivating the generalist policy based on different sensory feedback and commands at run time. We experimentally analyze and demonstrate the versatile locomotion delivered by GALA on both simulated and real quadruped robots in various environments. The code is available at https://github.com/vsislab/GALA.
Yanyun Chen, Ran Song 0001, Jiapeng Sheng, Wenhao Tan, Yibin Li 0001, Wei Zhang 0021
IEEE Trans Autom. Sci. Eng.7
2025 Autonomous and Adaptive Role Selection for Multi-Robot Collaborative Area Search Based on Deep Reinforcement Learning
abstract
In the tasks of multi-robot collaborative area search, we propose the unified approach for simultaneous mapping for sensing more targets (exploration) while searching and locating the targets (coverage). Specifically, we implement a hierarchical multi-agent reinforcement learning algorithm to decouple task planning from task execution. The role concept is integrated into the upper-level task planning for role selection, which enables robots to learn the role based on the state status from the upper-view. Besides, an intelligent role switching mechanism enables the role selection module to function between two timesteps, promoting both exploration and coverage interchangeably. Then, the primitive policy learns how to plan based on their assigned roles and local observation for sub-task execution. The well-designed experiments show the scalability and generalization of our method compared with state-of-the-art approaches in scenes with varying complexity and numbers of robots. Our code is released at https://github.com/linaug/Role_selection.
Jiyu Cheng, Hao Zhang 0113, Zhichao Cui, Wei Zhang 0021, Yuehu Liu
IEEE Trans Autom. Sci. Eng.5
2025 TRNet: Two-Tier Recursion Network for Co-Salient Object Detection
abstract
Co-salient object detection (CoSOD) is to find the salient and recurring objects from a series of relevant images, where modeling inter-image relationships plays a crucial role. Different from the commonly used direct learning structure that inputs all the intra-image features into some well-designed modules to represent the inter-image relationship, we resort to adopting a recursive structure for inter-image modeling, and propose a two-tier recursion network (TRNet) to achieve CoSOD in this paper. The two-tier recursive structure of the proposed TRNet is embodied in two stages of inter-image extraction and distribution. On the one hand, considering the task adaptability and inter-image correlation, we design an inter-image exploration with recursive reinforcement module to learn the local and global inter-image correspondences, guaranteeing the validity and discriminativeness of the information in the step-by-step propagation. On the other hand, we design a dynamic recursion distribution module to fully exploit the role of inter-image correspondences in a recursive structure, adaptively assigning common attributes to each individual image through an improved semi-dynamic convolution. Experimental results on five prevailing CoSOD benchmarks demonstrate that our TRNet outperforms other competitors in terms of various evaluation metrics. The code and results of our method are available athttps://github.com/rmcong/TRNet_TCSVT2025.
Runmin Cong, Ning Yang 0008, Hongyu Liu 0003, Dingwen Zhang, Qingming Huang, Sam Kwong, Wei Zhang 0021
IEEE Trans. Circuits Syst. Video Technol.7
2025 A Deep Reinforcement Learning Approach Using Asymmetric Self-Play for Robust Multirobot Flocking
abstract
Flocking control, as an essential approach for survivable navigation of multirobot systems, has been widely applied in fields, such as logistics, service delivery, and search and rescue. However, realistic environments are typically complex, dynamic, and even aggressive, posing considerable threats to the safety of flocking robots. In this article, based on deep reinforcement learning, anAsymmetricSelf-play-empoweredFlockingControl framework is proposed to address this concern. Specifically, the flocking robots are trained concurrently with learnable adversarial interferers to stimulate the intelligence of the flocking strategy. A two-stage self-play training paradigm is developed to improve the robustness and generalization of the model. Furthermore, an auxiliary training module regarding the learning of transition dynamics is designed, dramatically enhancing the adaptability to environmental uncertainties. Feature-level and agent-level attention are implemented for action and value generation, respectively. Both extensive comparative experiments and real-world deployment demonstrate the superiority and practicality of the proposed framework.
Yunjie Jia, Yong Song 0005, Jiyu Cheng, Jiong Jin, Wei Zhang 0021, Simon X. Yang, Sam Kwong
IEEE Trans. Ind. Informatics5
2025 Empowering Multirobot Flocking in Complex Environments via Effective Communication: A Deep Reinforcement Learning Approach
abstract
Multirobot flocking is crucial for safe and cooperative navigation, with wide applications in logistics, service delivery, and mobile surveillance. Despite significant progress, developing effective flocking strategies under complex conditions remains challenging. Communication is a vital technique for multirobot coordination. In this article, we propose refinement and enhancement of communication information (REIN), a novel deep reinforcement learning-based framework designed to improve communication effectiveness in leader–follower flocking systems through the REIN. First, regarding information refinement, a graph-based information refiner, integrating directed graph-structured communication with an innovative edge filter, is developed for selective multirobot interaction. It helps robots adaptively focus on relevant neighbors, considerably alleviating information overload. Second, for information enhancement, a cognition-aligned information enhancer is designed that boosts information expressiveness by encouraging team consensus. It utilizes two cascaded leader-related objectives to optimize information towards cognitive alignment among decentralized followers. Extensive comparisons with state-of-the-art approaches and ablation versions demonstrate the superiority of our framework. Physical experiments are also conducted to validate its practicality.
Yunjie Jia, Yong Song 0005, Jiyu Cheng, Heteng Zhang, Wei Zhang 0021, Rui Song 0002, Simon X. Yang, Sam Kwong
IEEE Trans. Ind. Informatics5
2025 Reference-Based Iterative Interaction With P2-Matching for Stereo Image Super-Resolution
abstract
Stereo Image Super-Resolution (SSR) holds great promise in improving the quality of stereo images by exploiting the complementary information between left and right views. Most SSR methods primarily focus on the inter-view correspondences in low-resolution (LR) space. The potential of referencing a high-quality SR image of one view benefits the SR for the other is often overlooked, while those with abundant textures contribute to accurate correspondences. Therefore, we propose Reference-based Iterative Interaction (RIISSR), which utilizes reference-based iterative pixel-wise and patch-wise matching, dubbed $P^{2}$ -Matching, to establish cross-view and cross-resolution correspondences for SSR. Specifically, we first design the information perception block (IPB) cascaded in parallel to extract hierarchical contextualized features for different views. Pixel-wise matching is embedded between two parallel IPBs to exploit cross-view interaction in LR space. Iterative patch-wise matching is then executed by utilizing the SR stereo pair as another mutual reference, capitalizing on the cross-scale patch recurrence property to learn high-resolution (HR) correspondences for SSR performance. Moreover, we introduce the supervised side-out modulator (SSOM) to re-weight local intra-view features and produce intermediate SR images, which seamlessly bridge two matching mechanisms. Experimental results demonstrate the superiority of RIISSR against existing state-of-the-art methods.
Runmin Cong, Rongxin Liao, Feng Li 0037, Ronghui Sheng, Huihui Bai 0001, Renjie Wan, Sam Kwong, Wei Zhang 0021
IEEE Trans. Image Process.8
2025 ViV-ReID: Bidirectional Structural-Aware Spatial-Temporal Graph Networks on Large-Scale Video-Based Vessel Re-Identification Dataset
abstract
Vessel re-identification (ReID) serves as a foundational task for intelligent maritime transportation systems. To enhance maritime surveillance capabilities, this study investigates video-based vessel ReID, a critical yet underexplored task in intelligent transportation systems. The lack of relevant datasets has limited the progress of Video-based vessel ReID research work. We established ViV-ReID, the first publicly available large-scale video-based vessel ReID dataset, comprising 480 vessel identities captured from 20 cross-port camera views (7,165 tracklets and 1.14 million frames), establishing a benchmark for advancing vessel ReID from image to video processing. Videos offer significantly richer information than single-frame images. The dynamic nature of video often leads to fragmented spatio-temporal features causing disrupted contextual understanding, and to address this problem, we further propose a Bidirectional Structural-Aware Spatial-Temporal Graph Network (Bi-SSTN) that explicitly aligns spatio-temporal features using vessel structural priors. Extensive experiments on the ViV-ReID dataset demonstrate that image-based ReID methods often show suboptimal performance when applied to video data. Meanwhile, it is crucial to validate the effectiveness of spatio-temporal information and establish performance benchmarks for different methods. The Bidirectional Structural-Aware Spatial-Temporal Graph Network (Bi-SSTN) significantly outperforms state-of-the-art methods on ViV-ReID, confirming its efficacy in modeling vessel-specific spatio-temporal patterns. Project web page: https://vsislab.github.io/ViV_ReID/.
Mingxin Zhang 0006, Fuxiang Feng, Lin Zhang 0041, Youmei Zhang, Xiaolei Li 0003, Wei Zhang 0021
IEEE Trans. Image Process.7
2025 Fast 3D Room Layout Estimation Based on Compact High-Level Representation
abstract
3D room layout estimation aims to reconstruct the holistic 3D structure from an indoor RGB image. For most of the deep learning-based methods, layout inference is guided by a kind of learned 2D mid-level representation such as pixel-wise surface labels. However, learning such high-resolution 2D representation might suffer from information redundancy and memory consumption, and will increase the runtime of estimation and deployment cost for practical applications. In this paper, we attempt to learn a compact high-level representation with only 29 real numbers for estimating the 3D layout using general regression networks. The learned compact high-level representation contains three components: instance-wise plane parameters, camera intrinsic parameters, and plane location indicators. With the learned representation, the inverse depth map of each plane can be calculated to reconstruct the 3D layout. We further design a set of order-agnostic loss functions to restrict the produced inverse depth maps, with which the model can be trained with either weak 2D layout labels or full 3D layout supervision. Moreover, by jointly learning the plane parameters and locations, the model is benefited from 3D reasoning. Experimental results show that our method is much faster than the existing layout estimation methods and obtains competitive performance on benchmark datasets, showing its potential for real-time applications.
Weidong Zhang 0005, Yu Qiao 0001, Ying Liu 0026, Ran Song 0001, Wei Zhang 0021
IEEE Trans. Image Process.5
2025 SLPDR: A Benchmark for Ship License Plate Detection and Recognition
abstract
Ship identification is a prerequisite for the intelligent management of maritime transportation, yet existing research is confined to broad ship detection and categorization, which only provides the ship’s location or type instead of its identification. Inspired by the research on the Car License Plate (CLP), we make the first attempt to propose the concept of the Ship License Plate (SLP). In addition, the limited data hinders research on ship identification. To overcome this obstacle, we construct the first large-scale Ship License Plate Detection and Recognition (SLPDR) dataset, which contains 1,472 ship identities and 88,862 images. In addition, this paper proposes an SLP detection model named YOLO-SSA and evaluates this model as well as typical detection methods on the SLPDR dataset. The experimental results demonstrate that the proposed YOLO-SSA achieves better SLP detection performance by enhancing the features where ships and SLPs are located. Furthermore, we explore the prospective applications of SLPs in intelligent maritime transportation, including ship monitoring and berth management. Project web page: https://vsislab.github.io/SLPDR/
Youmei Zhang, Ran Song 0001, Yonghuai Liu, Ardhendu Behera, Mingxin Zhang 0006, Wei Zhang 0021
IEEE Trans. Intell. Transp. Syst.7
2025 Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement Learning
abstract
Despite significant advancements in multirobot technologies, efficiently and collaboratively exploring an unknown environment remains a major challenge. In this paper, we propose AIM-Mapping, an Asymmetric InforMation enhanced Mapping framework based on deep reinforcement learning. The framework fully leverages the privileged information to help construct the environmental representation as well as the supervised signal in an asymmetric actor-critic training framework. Specifically, privileged information is used to evaluate exploration performance through an asymmetric feature representation module and a mutual information evaluation module. The decision-making network employs the trained feature encoder to extract structural information of the environment and integrates it with a topological map constructed based on geometric distance. By leveraging this topological map representation, we apply topological graph matching to assign corresponding boundary points to each robot as long-term goal points. We conduct experiments in both iGibson simulation environments and real-world scenarios. The results demonstrate that the proposed method achieves significant performance improvements compared to existing approaches.
Jiyu Cheng, Junhui Fan, Xiaolei Li 0003, Paul L. Rosin, Yibin Li 0001, Wei Zhang 0021
IEEE Trans. Robotics6
2024 DDLNet: Boosting Remote Sensing Change Detection with Dual-Domain Learning
abstract
Remote sensing change detection (RSCD) aims to identify the changes of interest in a region by analyzing multi-temporal remote sensing images, and has an outstanding value for local development monitoring. Existing RSCD methods are devoted to contextual modeling in the spatial domain to enhance the changes of interest. Despite the satisfactory performance achieved, the lack of knowledge in the frequency domain limits the further improvement of model performance. In this paper, we propose DDLNet, a RSCD network based on dual-domain learning (i.e., frequency and spatial domains). In particular, we design a Frequency-domain Enhancement Module (FEM) to capture frequency components from the input bi-temporal images using Discrete Cosine Transform (DCT) and thus enhance the changes of interest. Besides, we devise a Spatial-domain Recovery Module (SRM) to fuse spatiotemporal features for reconstructing spatial details of change representations. Extensive experiments on three benchmark RSCD datasets demonstrate that the proposed method achieves state-of-the-art performance and reaches a more satisfactory accuracy-efficiency trade-off. Our code is publicly available at https://github.com/xwmaxwma/rschange.
Rui Che, Huanting Zhang, Wei Zhang 0021
ICME5
2024 ESNet: Evolution and Succession Network for High-Resolution Salient Object Detection
abstract
Preserving details and avoiding high computational costs are the two main challenges for the High-Resolution Salient Object Detection (HRSOD) task. In this paper, we propose a two-stage HRSOD model from the perspective of evolution and succession, including an evolution stage with Low-resolution Location Model (LrLM) and a succession stage with High-resolution Refinement Model (HrRM). The evolution stage achieves detail-preserving salient objects localization on the low-resolution image through the evolution mechanisms on supervision and feature; the succession stage utilizes the shallow high-resolution features to complement and enhance the features inherited from the first stage in a lightweight manner and generate the final high-resolution saliency prediction. Besides, a new metric named Boundary-Detail-aware Mean Absolute Error (${MAE}_{{BD}}$) is designed to evaluate the ability to detect details in high-resolution scenes. Extensive experiments on five datasets demonstrate that our network achieves superior performance at real-time speed (49 FPS) compared to state-of-the-art methods.
Hongyu Liu 0003, Runmin Cong, Hua Li 0012, Qianqian Xu 0001, Qingming Huang, Wei Zhang 0021
ICML6
2024 MPP: Multiscale Path Planning for UGV Navigation in Semi-structured Environments
abstract
Autonomous navigation of unmanned ground vehicles (UGVs) in structured road and indoor environments has made significant progress in recent years. However, navigation in outdoor semi-structured environments remains a challenge. This paper presents the multiscale path planning (MPP) method for UGV navigation in semi-structured environments. MPP leverages global, mid-layer and local planners to obtain global path and handle local obstacles of different sizes. First, the global planner provides guidance based on road connection relationships, selecting optimal connections by evaluating the distance between road nodes. Next, the mid-layer planner perceives large-scale obstacles and constructs the costmap, generating a mid-layer path that offers a general direction for the UGV. Finally, a local trajectory planning algorithm, namely terrain-considering timed elastic band (TC-TEB), is used to obtain local trajectory. This algorithm incorporates terrain-velocity constraints into the TEB algorithm to ensure the vehicle’s vertical stability. We demonstrate the safety and effectiveness of MPP through experiments in both simulated and real-world environments.
Ran Song 0001, Wei Zhang 0021
IROS6
2024 MPGNet: Learning Move-Push-Grasping Synergy for Target-Oriented Grasping in Occluded Scenes
abstract
This paper focuses on target-oriented grasping in occluded scenes, where the target object is specified by a binary mask and the goal is to grasp the target object with as few robotic manipulations as possible. Most existing methods rely on a push-grasping synergy to complete this task. To deliver a more powerful target-oriented grasping pipeline, we present MPGNet, a three-branch network for learning a synergy between moving, pushing, and grasping actions. We also propose a multi-stage training strategy to train the MPGNet which contains three policy networks corresponding to the three actions. The effectiveness of our method is demonstrated via both simulated and real-world experiments. Video of the real-world experiments is at https://youtu.be/S_QKZqkh0w8.
Dayou Li, Chenkun Zhao, Ran Song 0001, Xiaolei Li 0003, Wei Zhang 0021
IROS6
2024 Coarse-to-Fine Detection of Multiple Seams for Robotic Welding
abstract
Efficiently detecting target weld seams while ensuring sub-millimeter accuracy has always been an important challenge in autonomous welding, which has significant application in industrial practice. Previous works mostly focused on recognizing and localizing welding seams one by one, leading to inferior efficiency in modeling the workpiece. This paper proposes a novel framework capable of multiple weld seams extraction using both RGB images and 3D point clouds. The RGB image is used to obtain the region of interest by approximately localizing the weld seams, and the point cloud is used to achieve the fine-edge extraction of the weld seams within the region of interest using region growth. Our method is further accelerated by using a pre-trained deep learning model to ensure both efficiency and generalization ability. The proposed method was comprehensively tested on various workpieces featuring both linear and curved weld seams, as well as in physical experiment systems. The results showcase considerable potential for real-world industrial applications, emphasizing the method’s efficiency and effectiveness. Videos of the real-world experiments can be found at https://youtu.be/pq162HSP2D4.
Pengkun Wei, Dayou Li, Ran Song 0001, Wei Zhang 0021
IROS6
2024 Exploiting Inter-Sample Affinity for Knowability-Aware Universal Domain Adaptation
Yifan Wang 0020, Lin Zhang 0041, Ran Song 0001, Hongliang Li 0001, Paul L. Rosin, Wei Zhang 0021
Int. J. Comput. Vis.6
2024 DeTAL: Open-Vocabulary Temporal Action Localization With Decoupled Networks
abstract
Pre-trained visual-language (ViL) models have demonstrated good zero-shot capability in video understanding tasks, where they were usually adapted through fine-tuning or temporal modeling. However, in the task of open-vocabulary temporal action localization (OV-TAL), such adaption reduces the robustness of ViL models against different data distributions, leading to a misalignment between visual representations and text descriptions of unseen action categories. As a result, existing methods often strike a trade-off between action detection and classification. Aiming at this issue, this paper proposes DeTAL, a simple but effective two-stage approach for OV-TAL. DeTAL decouples action detection from action classification to avoid the compromise between them, and the state-of-the-art methods for close-set action localization can be handily adapted to OV-TAL, which significantly improves the performance. Meanwhile, DeTAL can easily tackle the scenario where action category annotations are unavailable in the training dataset. In the experiments, we propose a new cross-dataset setting to evaluate the zero-shot capability of different methods. And the results demonstrate that DeTAL outperforms the state-of-the-art methods for OV-TAL on both THUMOS14 and ActivityNet1.3.
Zhiheng Li 0005, Ran Song 0001, Lin Ma 0002, Wei Zhang 0021
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 IGCN: A Provably Informative GCN Embedding for Semi-Supervised Learning With Extremely Limited Labels
abstract
Graph Neural Networks (GNNs) have gained much more attention in the representation learning for the graph-structured data. However, the labels are always limited in the graph, which easily leads to the overfitting problem and causes the poor performance. To solve this problem, we propose a new framework called IGCN, short for Informative Graph Convolutional Network, where the objective of IGCN is designed to obtain the informative embeddings via discarding the task-irrelevant information of the graph data based on the mutual information. As the mutual information for irregular data is intractable to compute, our framework is optimized via a surrogate objective, where two terms are derived to approximate the original objective. For the former term, it demonstrates that the mutual information between the learned embeddings and the ground truth should be high, where we utilize the semi-supervised classification loss and the prototype based supervised contrastive learning loss for optimizing it. For the latter term, it requires that the mutual information between the learned node embeddings and the initial embeddings should be high and we propose to minimize the reconstruction loss between them to achieve the goal of maximizing the latter term from the feature level and the layer level, which contains the graph encoder-decoder module and a novel architecture GCN$_{Info}$. Moreover, we provably show that the designed GCN$_{Info}$can better alleviate the information loss and preserve as much useful information of the initial embeddings as possible. Experimental results show that the IGCN outperforms the state-of-the-art methods on 7 popular datasets.
Lin Zhang 0041, Ran Song 0001, Wenhao Tan, Lin Ma 0002, Wei Zhang 0021
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Prototype learning for adversarial domain adaptation
Yuchun Fang, Chen Chen 0114, Wei Zhang 0021, Zhaoxiang Zhang 0001, Shaorong Xie
Pattern Recognit.3
2024 A Hierarchical Framework for Quadruped Omnidirectional Locomotion Based on Reinforcement Learning
abstract
Quadruped locomotion is challenging for many learning-based algorithms. This is because it requires tedious manual tuning to cope with different types of terrains and is difficult to deploy in reality due to the sim-to-real gap between the training and the testing scenarios. This paper proposes a quadruped robot learning system for agile locomotion which does not require any pre-training and works well in various terrains. We introduce a hierarchical framework that uses reinforcement learning as the high-level policy to adjust the low-level trajectory generator for a better adaptability to various terrains. We compact the observation and the action spaces of reinforcement learning to deploy the proposed framework on a host computer interfaced with the robot. Besides, we design an omnidirectional trajectory generator guided by robot posture, which generates omnidirectional foot trajectories to interact with the environment. Experimental results and the supplementary video demonstrate that our hierarchical framework only trained in simulation can be easily deployed in the real world, and also has the advantages of fast convergence and good terrain adaptability.Note to Practitioners—This paper presents a hierarchical framework for quadruped robots. It combines a high-level reinforcement learning controller with a posture-guided trajectory generator to adaptively generate omnidirectional motions. Our method is easy to train as it converges fast and does not need to adjust a dozen or so of rewards. The quadruped robot can be deployed in a real environment directly after being trained in simulation. With the trained hierarchical framework deployed on a remote host computer, the robot works well in a variety of real-world environments unseen in the simulation.
Wenhao Tan, Wei Zhang 0021, Ran Song 0001, Yu Zheng 0001, Yibin Li 0001
IEEE Trans Autom. Sci. Eng.3
2024 Heuristics Integrated Deep Reinforcement Learning for Online 3D Bin Packing
abstract
Online 3D Bin Packing Problem (3D-BPP) has a wide range of industrial applications and there is an emerging research interest in learning optimal bin packing policy and deploying it for real logistics applications. From the heuristic methods to the deep reinforcement learning (DRL) methods, the previous works have proposed many solutions to solve the online 3D-BPP. However, none of them have studied what and how heuristics can be modelled into DRL to build a more effective and practical bin packing pipeline. In this work, we thoroughly investigate what heuristics can be used in online 3D-BPP and how to effectively integrate the heuristics with the DRL. First, we design 3 different heuristics based on the physical rules of the real world and the experiences of the human packers, including the Physics-Heuristics, the Packing-Heuristics and the Unpacking-Heuristics. Second, we model the 3 types of heuristics into the DRL framework and propose a novel heuristic DRL method to solve the online 3D-BPP. Extensive experimental results show that our method achieves state-of-the-art bin packing performance and the resulting real-world system is able to reliably finish the bin packing task in real logistics scenarios. Supplementary video is available athttps://www.youtube.com/watch?v=x8GpmEELq18. Note to Practitioners—The rapid growth of e-commerce has significantly increased the burden of human packers in logistic warehouses, where the workers need to pick the products from a conveyor and pack them into bins (i.e. the online 3D bin packing). Thus it is of great importance to develop intelligent robotic systems to replace human labor, which is a long-standing topic in the field of control and automation science. This paper makes a substantial contribution to the related field by studying the online 3D bin packing in terms of both the theory and practice. On the one hand, the simulated experiments suggest that the presented algorithm significantly improves the space utilization of bin packing. On the other hand, the robotic system developed based on the proposed method can favourably finish the bin packing task in real logistics scenarios, demonstrating the practical use of our approach. Consequently, the approach proposed in this paper is totally applicable in logistic warehouses and is promising to drastically improve the working efficiency of the product packing in real warehouses. In the future, we will extend the presented approach to pack irregular-shaped objects and then facilitate more logistics applications.
Shuai Song, Shilei Chu, Ran Song 0001, Jiyu Cheng, Yibin Li 0001, Wei Zhang 0021
IEEE Trans Autom. Sci. Eng.7
2024 HARDer-Net: Hardness-Guided Discrimination Network for 3D Early Activity Prediction
abstract
To predict the class label from a partially observable activity sequence can be quite challenging due to the high degree of similarity existing in early segments of different activities. In this paper, an innovative HARDness-Guided Discrimination Network (HARDer-Net) is proposed to evaluate the relationship between similar activity pairs that are extremely hard to discriminate. To train our HARDer-Net, an innovative adversarial learning scheme has been designed, providing our network with the strength to extract subtle discrimination information for the prediction of 3D early activities. Moreover, to enhance the adversarial learning scheme efficacy of our model for 3D early action prediction, we construct a Hardness-Guided bank that dynamically records the hard similar samples and conducts reward-guided selections of these recorded hard samples using a deep reinforcement learning scheme. The proposed method significantly enhances the capability of the model to discern fine-grained differences in early activity sequences. Several widely-used activity datasets are used to evaluate our proposed HARDer-Net, and we achieve state-of-the-art performance across all the evaluated datasets.
Wei Zhang 0021, Ling-Yu Duan, Jun Liu 0036
IEEE Trans. Circuits Syst. Video Technol.3
2024 Structural Digital Twin Modeling and Adaptive Pretrain-Finetune Learning for Dynamic Impact Identification on Wind Turbine Blades
abstract
Identification of dynamic impacts on wind turbine blades (WTB) is critical for structural safety and predictive maintenance. Although WTBs are exposed to impact risk, data samples of transient impact responses are rare in practice. To accurately identify dynamic impacts using limited real collections, this article proposes a systematic methodology with structural digital twin (SDT) modeling and adaptive pretrain-finetune learning. Given the generated source data from the SDT, deep transfer learning is addressed for the few-shot target data from real impacts. Particularly, receptive attention mechanisms are designed in an adaptively pretraining neural network for learning the SDT data in the source domain. In addition, a self-attention switch network is proposed as an adaptively finetuning neural network for knowledge transferring to the target domain. The physical experiments on both comparison and ablation study demonstrate the effectiveness and accuracy of the proposed method. Meanwhile, the identification process achieves high computational efficiency for online implementation.
Teng Li 0005, Yingxin Luan, Zhendong Pang, Wei Zhang 0021
IEEE Trans. Ind. Informatics4
2024 SRNSD: Structure-Regularized Night-Time Self-Supervised Monocular Depth Estimation for Outdoor Scenes
abstract
Deep CNNs have achieved impressive improvements for night-time self-supervised depth estimation form a monocular image. However, the performance degrades considerably compared to day-time depth estimation due to significant domain gaps, low visibility, and varying illuminations between day and night images. To address these challenges, we propose a novel night-time self-supervised monocular depth estimation framework with structure regularization, i.e., SRNSD, which incorporates three aspects of constraints for better performance, including feature and depth domain adaptation, image perspective constraint, and cropped multi-scale consistency loss. Specifically, we utilize adaptations of both feature and depth output spaces for better night-time feature extraction and depth map prediction, along with high- and low-frequency decoupling operations for better depth structure and texture recovery. Meanwhile, we employ an image perspective constraint to enhance the smoothness and obtain better depth maps in areas where the luminosity jumps change. Furthermore, we introduce a simple yet effective cropped multi-scale consistency loss that utilizes consistency among different scales of depth outputs for further optimization, refining the detailed textures and structures of predicted depth. Experimental results on different benchmarks with depth ranges of 40m and 60m, including Oxford RobotCar dataset, nuScenes dataset and CARLA-EPE dataset, demonstrate the superiority of our approach over state-of-the-art night-time self-supervised depth estimation approaches across multiple metrics, proving our effectiveness.
Runmin Cong, Chunlei Wu, Xibin Song, Wei Zhang 0021, Sam Kwong, Hongdong Li, Pan Ji
IEEE Trans. Image Process.4
2024 Learning Common Semantics via Optimal Transport for Contrastive Multi-View Clustering
abstract
Multi-view clustering aims to learn discriminative representations from multi-view data. Although existing methods show impressive performance by leveraging contrastive learning to tackle the representation gap between every two views, they share the common limitation of not performing semantic alignment from a global perspective, resulting in the undermining of semantic patterns in multi-view data. This paper presents CSOT, namely Common Semantics via Optimal Transport, to boost contrastive multi-view clustering via semantic learning in a common space that integrates all views. Through optimal transport, the samples in multiple views are mapped to the joint clusters which represent the multi-view semantic patterns in the common space. With the semantic assignment derived from the optimal transport plan, we design a semantic learning module where the soft assignment vector works as a global supervision to enforce the model to learn consistent semantics among all views. Moreover, we propose a semantic-aware re-weighting strategy to treat samples differently according to their semantic significance, which improves the effectiveness of cross-view contrastive representation learning. Extensive experimental results demonstrate that CSOT achieves the state-of-the-art clustering performance.
Qian Zhang 0076, Lin Zhang 0041, Ran Song 0001, Runmin Cong, Yonghuai Liu, Wei Zhang 0021
IEEE Trans. Image Process.6
2024 Hierarchical Perception-Improving for Decentralized Multi-Robot Motion Planning in Complex Scenarios
abstract
Multi-robot cooperative navigation is an important task, which has been widely studied in many fields like logistics, transportation, and disaster rescue. However, most of the existing methods either require some strong assumptions or are validated in simple scenarios, which greatly hinders their implementation in the real world. In this paper, more complex environments are considered in which robots can only acquire local observations from their own sensors and have only limited communication capabilities for mapless collaborative navigation. To address this challenging task, we propose a hierarchical framework, by fusing bothSensor-wise andAgent-wise features forPerception-Improving (SAPI), which can adaptively integrate features from different information sources to improve perception capabilities. Specifically, to facilitate scene understanding, we assign prior knowledge to the visual coder to generate efficient embeddings. For effective feature representation, an attention-based sensor fusion network is designed to fuse sensor-level information of visual and LiDAR sensors, while graph convolution with multi-head attention mechanism is applied to aggregate agent-level information from an arbitrary number of neighbors. In addition, reinforcement learning is used to optimize the policy, where a novel compound reward function is introduced to guide training. Extensive experiments demonstrate that our method has excellent generalization ability in different scenarios and scalability for large-scale systems.
Yunjie Jia, Yong Song 0005, Bo Xiong 0001, Jiyu Cheng, Wei Zhang 0021, Simon X. Yang, Sam Kwong
IEEE Trans. Intell. Transp. Syst.5
2024 Ship Landmark: An Informative Ship Image Annotation and Its Applications
abstract
Visual perception of ships has been attracting increasing attention in the fields of computer vision and ocean engineering. Despite the extensive work related to landmark detection of common objects, the role of landmarks in ship perception has been overlooked. In this paper, we aim to fill this gap by focusing on ship landmarks. Specifically, we give a comprehensive analysis of both the physical structure and deep features of ships, which finds that highlighted areas in feature maps correspond with structurally significant parts of ships. By summarizing the locations of such areas in ships, we define 20 ship landmarks and build the Ship Landmark Dataset (SLAD), the first ship dataset with landmark annotations. We also provide a benchmark for ship landmark detection by evaluating state-of-the-art landmark detection methods on the newly built SLAD. Moreover, we showcased several applications of ship landmarks, including ship recognition, ship image generation, key area detection for ships, and ship detection. Project web page:https://vsislab.github.io/Ships_VSIS/.
Mingxin Zhang 0006, Qian Zhang 0076, Ran Song 0001, Paul L. Rosin, Wei Zhang 0021
IEEE Trans. Intell. Transp. Syst.5
2024 Multi-Robot Environmental Coverage With a Two-Stage Coordination Strategy via Deep Reinforcement Learning
abstract
Multi-robot environmental coverage can be widely used in many applications like search and rescue. However, it is challenging to coordinate the robot team for high coverage efficiency. In this paper, we propose a Two-Stage Coordination (TSC) strategy, which consists of a high-level leader module and a low-level action executor. The former provides the robots with the topology and geometry of the environment, which are crucial for robots to learn “where” they should go and avoid invalid coverage. Based on the observed information and the environmental topology, the latter module takes primitive action to reach the sub-goal. To facilitate cooperation among the robots, we aggregate local perception information of neighbors from different hops based on graph neural networks. We compare our method with state-of-the-art multi-robot coverage approaches. Experiments and supporting ablation studies show the superior efficiency, scalability, and generalization of our algorithm especially in unseen style and scale of scenes, and an unseen number of robots.
Jiyu Cheng, Hao Zhang 0113, Wei Zhang 0021, Yuehu Liu
IEEE Trans. Intell. Transp. Syst.4
2024 Query-Guided Prototype Evolution Network for Few-Shot Segmentation
abstract
Previous Few-Shot Segmentation (FSS) approaches exclusively utilize support features for prototype generation, neglecting the specific requirements of the query. To address this, we present the Query-guided Prototype Evolution Network (QPENet), a new method that integrates query features into the generation process of foreground and background prototypes, thereby yielding customized prototypes attuned to specific queries. The evolution of the foreground prototype is accomplished through a support-query-support iterative process involving two new modules: Pseudo-prototype Generation (PPG) and Dual Prototype Evolution (DPE). The PPG module employs support features to create an initial prototype for the preliminary segmentation of the query image, resulting in a pseudo-prototype reflecting the unique needs of the current query. Subsequently, the DPE module performs reverse segmentation on support images using this pseudo-prototype, leading to the generation of evolved prototypes, which can be considered as custom solutions. As for the background prototype, the evolution begins with a global background prototype that represents the generalized features of all training images. We also design a Global Background Cleansing (GBC) module to eliminate potential adverse components mirroring the characteristics of the current foreground class. Experimental results on the PASCAL-52and COCO-202datasets attest to the substantial enhancements achieved by QPENet over prevailing state-of-the-art techniques, underscoring the validity of our ideas.
Runmin Cong, Jinpeng Chen 0003, Wei Zhang 0021, Qingming Huang, Yao Zhao 0001
IEEE Trans. Multim.4
2024 Neighborhood-Aware Mutual Information Maximization for Source-Free Domain Adaptation
abstract
Recently, the source-free domain adaptation (SFDA) problem has attracted much attention, where the pre-trained model for the source domain is adapted to the target domain in the absence of source data. However, due to domain shift, the negative alignment usually exists between samples from the same class, which may lower intra-class feature similarity. To address this issue, we present a self-supervised representation learning strategy for SFDA, named as neighborhood-aware mutual information (NAMI), which maximizes the mutual information (MI) between the representations of target samples and their corresponding neighbors. Moreover, we theoretically demonstrate that NAMI can be decomposed into a weighted sum of local MI, which suggests that the weighted terms can better estimate NAMI. To this end, we introduce neighborhood consensus score over the set of weakly and strongly augmented views and point-wise density based on neighborhood, both of which determine the weights of local MI for NAMI by leveraging the neighborhood information of samples. The proposed method can significantly handle domain shift and adaptively reduce the noise in the neighborhood of each target sample. In combination with the consistency loss over views, NAMI leads to consistent improvement over existing state-of-the-art methods on three popular SFDA benchmarks.
Lin Zhang 0041, Yifan Wang 0020, Ran Song 0001, Mingxin Zhang 0006, Xiaolei Li 0003, Wei Zhang 0021
IEEE Trans. Multim.6
2023 E2E-LOAD: End-to-End Long-form Online Action Detection
abstract
Recently, feature-based methods for Online Action Detection (OAD) have been gaining traction. However, these methods are constrained by their fixed backbone design, which fails to leverage the potential benefits of a trainable backbone. This paper introduces an end-to-end learning network that revises these approaches, incorporating a backbone network design that improves effectiveness and efficiency. Our proposed model utilizes a shared initial spatial model for all frames and maintains an extended sequence cache, which enables low-cost inference. We promote an asymmetric spatiotemporal model that caters to long-form and short-form modeling. Additionally, we propose an innovative and efficient inference mechanism that accelerates extensive spatiotemporal exploration. Through comprehensive ablation studies and experiments, we validate the performance and efficiency of our proposed method. Remarkably, we achieve an end-to-end learning OAD of 17.3 (+12.6) FPS with 72.4% (+1.2%), 90.3% (+0.7%), and 48.1% (+26.0%) mAP on THMOUS’14, TVSeries, and HDD, respectively. The source code is available at https://github.com/sqiangcao99/E2E-LOAD.
Shuqiang Cao, Weixin Luo, Bairui Wang, Wei Zhang 0021, Lin Ma 0002
ICCV4
2023 WaterMask: Instance Segmentation for Underwater Imagery
abstract
Underwater image instance segmentation is a fundamental and critical step in underwater image analysis and understanding. However, the paucity of general multiclass instance segmentation datasets has impeded the development of instance segmentation studies for underwater images. In this paper, we propose the first underwater image instance segmentation dataset (UIIS), which provides 4628 images for 7 categories with pixel-level annotations. Meanwhile, we also design WaterMask for underwater image instance segmentation for the first time. In Water-Mask, we first devise Difference Similarity Graph Attention Module (DSGAT) to recover lost detailed information due to image quality degradation and downsampling to help the network prediction. Then, we propose Multi-level Feature Refinement Module (MFRM) to predict foreground masks and boundary masks separately by features at different scales, and guide the network through Boundary Mask Strategy (BMS) with boundary learning loss to provide finer prediction results. Extensive experimental results demonstrates that WaterMask can achieve significant gains of 2.9, 3.8 mAP over Mask R-CNN when using ResNet-50 and ResNet-101. Code and Dataset are available at https://github.com/LiamLian0727/WaterMask.
Shijie Lian, Hua Li 0012, Runmin Cong, Suqi Li, Wei Zhang 0021, Sam Kwong
ICCV5
2023 SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow Detection
abstract
Despite significant progress in shadow detection, current methods still struggle with the adverse impact of background color, which may lead to errors when shadows are present on complex backgrounds. Drawing inspiration from the human visual system, we treat the input shadow image as a composition of a background layer and a shadow layer, and design a Style-guided Dual-layer Disentanglement Network (SDDNet) to model these layers independently. To achieve this, we devise a Feature Separation and Recombination (FSR) module that decomposes multi-level features into shadow-related and background-related components by offering specialized supervision for each component, while preserving information integrity and avoiding redundancy through the reconstruction constraint. Moreover, we propose a Shadow Style Filter (SSF) module to guide the feature disentanglement by focusing on style differentiation and uniformization. With these two modules and our overall pipeline, our model effectively minimizes the detrimental effects of background color, yielding superior performance on three public datasets with a real-time inference speed of 32 FPS. Our code is publicly available at:https://github.com/rmcong/SDDNet_ACMMM23.
Runmin Cong, Yuchen Guan, Jinpeng Chen 0003, Wei Zhang 0021, Yao Zhao 0001, Sam Kwong
ACM Multimedia4
2023 Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection
abstract
By integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent years, the important role of Convolutional Neural Networks (CNNs) in feature extraction and cross-modality interaction has been fully explored, but it is still insufficient in modeling global long-range dependencies of self-modality and cross-modality. To this end, we introduce CNNs-assisted Transformer architecture and propose a novel RGB-D SOD network with Point-aware Interaction and CNN-induced Refinement (PICR-Net). On the one hand, considering the prior correlation between RGB modality and depth modality, an attention-triggered cross-modality point-aware interaction (CmPI) module is designed to explore the feature interaction of different modalities with positional constraints. On the other hand, in order to alleviate the block effect and detail destruction problems brought by the Transformer naturally, we design a CNN-induced refinement (CNNR) unit for content refinement and supplementation. Extensive experiments on five RGB-D SOD datasets show that the proposed network achieves competitive results in both quantitative and qualitative comparisons. Our code is publicly available at: https://github.com/rmcong/PICR-Net_ACMMM23.
Runmin Cong, Hongyu Liu 0003, Chen Zhang 0013, Wei Zhang 0021, Feng Zheng 0001, Ran Song 0001, Sam Kwong
ACM Multimedia4
2023 Frequency Perception Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to accurately detect objects hidden in the surrounding environment. However,the existing COD methods mainly locate camouflaged objects in the RGB domain, their performance has not been fully exploited in many challenging scenarios. Considering that the features of the camouflaged object and the background are more discriminative in the frequency domain, we propose a novel learnable and separable frequency perception mechanism driven by the semantic hierarchy in the frequency domain. Our entire network adopts a two-stage model, including a frequency-guided coarse localization stage and a detail-preserving fine localization stage.With the multi-level features extracted by the backbone, we design a flexible frequency perception module based on octave convolution for coarse positioning. Then, we design the correction fusion module to step-by-step integrate the high-level features through the prior-guided correction and cross-layer feature channel association, and finally combine them with the shallow features to achieve the detailed correction of the camouflaged objects. Compared with the currently existing models, our proposed method achieves competitive performance in three popular benchmark datasets both qualitatively and quantitatively. The code will be released at https://github.com/rmcong/FPNet_ACMMM23.
Runmin Cong, Mengyao Sun 0003, Sanyi Zhang, Xiaofei Zhou 0003, Wei Zhang 0021, Yao Zhao 0001
ACM Multimedia5
2023 3D Visual Saliency: An Independent Perceptual Measure or a Derivative of 2D Image Saliency?
abstract
While 3D visual saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and has been well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art 3D visual saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that 3D visual saliency might associate with 2D image saliency. This paper proposes a framework that combines a Generative Adversarial Network and a Conditional Random Field for learning visual saliency of both a single 3D object and a scene composed of multiple 3D objects with image saliency ground truth to 1) investigate whether 3D visual saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting 3D visual saliency. Through extensive experiments, we not only demonstrate that our method significantly outperforms the state-of-the-art approaches, but also manage to answer the interesting and worthy question proposed within the title of this paper.
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 PUGAN: Physical Model-Guided Underwater Image Enhancement Using GAN With Dual-Discriminators
abstract
Due to the light absorption and scattering induced by the water medium, underwater images usually suffer from some degradation problems, such as low contrast, color distortion, and blurring details, which aggravate the difficulty of downstream underwater understanding tasks. Therefore, how to obtain clear and visually pleasant images has become a common concern of people, and the task of underwater image enhancement (UIE) has also emerged as the times require. Among existing UIE methods, Generative Adversarial Networks (GANs) based methods perform well in visual aesthetics, while the physical model-based methods have better scene adaptability. Inheriting the advantages of the above two types of models, we propose a physical model-guided GAN model for UIE in this paper, referred to as PUGAN. The entire network is under the GAN architecture. On the one hand, we design a Parameters Estimation subnetwork (Par-subnet) to learn the parameters for physical model inversion, and use the generated color enhancement image as auxiliary information for the Two-Stream Interaction Enhancement sub-network (TSIE-subnet). Meanwhile, we design a Degradation Quantization (DQ) module in TSIE-subnet to quantize scene degradation, thereby achieving reinforcing enhancement of key regions. On the other hand, we design the Dual-Discriminators for the style-content adversarial constraint, promoting the authenticity and visual aesthetics of the results. Extensive experiments on three benchmark datasets demonstrate that our PUGAN outperforms state-of-the-art methods in both qualitative and quantitative metrics. The code and results can be found from the link of https://rmcong.github.io/proj_PUGAN.html.
Runmin Cong, Wei Zhang 0021, Chongyi Li, Chunle Guo, Qingming Huang, Sam Kwong
IEEE Trans. Image Process.3
2023 SMAM: Self and Mutual Adaptive Matching for Skeleton-Based Few-Shot Action Recognition
abstract
This paper focuses on skeleton-based few-shot action recognition. Since skeleton is essentially a sparse representation of human action, the feature maps extracted from it, through a standard encoder network in the few-shot condition, may not be sufficiently discriminative for some action sequences that look partially similar to each other. To address this issue, we propose a self and mutual adaptive matching (SMAM) module to convert such feature maps into more discriminative feature vectors. Our method, named as SMAM-Net, first leverages both the temporal information associated with each individual skeleton joint and the spatial relationship among them for feature extraction. Then, the SMAM module adaptively measures the similarity between labeled and query samples and further carries out feature matching within the query set to distinguish similar skeletons of various action categories. Experimental results show that the SMAM-Net outperforms other baselines on the large-scale NTU RGB + D 120 dataset in the tasks of one-shot and five-shot action recognition. We also report our results on smaller datasets including NTU RGB + D 60, SYSU and PKU-MMD to demonstrate that our method is reliable and generalises well on different datasets. Codes and the pretrained SMAM-Net will be made publicly available.
Zhiheng Li 0005, Xuyuan Gong, Ran Song 0001, Peng Duan 0002, Jun Liu 0036, Wei Zhang 0021
IEEE Trans. Image Process.6
2023 Unsupervised Maritime Vessel Re-Identification With Multi-Level Contrastive Learning
abstract
Re-identification (re-ID) of maritime vessels plays an important role in marine surveillance, but remains highly unexplored due to the lack of large-scale annotated datasets. In vessel re-ID, contrastive methods are supposed to learn discriminative representation from unlabeled vessel images in an unsupervised manner. However, directly introducing classical instance-level contrastive methods to maritime vessel re-ID suffers from the difficulty of finding vessel images with the same pseudo label as positive images, which potentially leads to inefficient training and unsatisfactory performance. This paper proposes a simple but effective method to solve such a hard positive problem. Our method takes all images in an intra-batch cluster as positives and excludes them from the set of negative samples when computing instance-level contrastive loss. Based on this strategy, we construct a multi-level contrastive learning (MCL) framework for vessel re-ID trained with the specifically designed intra-batch cluster-level contrastive loss along with the instance-level one. Experiments on a newly proposed dataset consisting of 1,248 vessel identities show that MCL achieves the state-of-the-art performance compared with other unsupervised methods.
Qian Zhang 0076, Mingxin Zhang 0006, Jinghe Liu, Xuanyu He, Ran Song 0001, Wei Zhang 0021
IEEE Trans. Intell. Transp. Syst.6
2023 Classification of Brain Disorders in rs-fMRI via Local-to-Global Graph Neural Networks
abstract
Recently, functional brain network has been used for the classification of brain disorders, such as Autism Spectrum Disorder (ASD) and Alzheimer's disease (AD). Existing methods either ignore the non-imaging information associated with the subjects and the relationship between the subjects, or cannot identify and analyze disease-related local brain regions and biomarkers, leading to inaccurate classification results. This paper proposes a local-to-global graph neural network (LG-GNN) to address this issue. A local ROI-GNN is designed to learn feature embeddings of local brain regions and identify biomarkers, and a global Subject-GNN is then established to learn the relationship between the subjects with the embeddings generated by the local ROI-GNN and the non-imaging information. The local ROI-GNN contains a self-attention based pooling module to preserve the embeddings most important for the classification. The global Subject-GNN contains an adaptive weight aggregation block to generate the multi-scale feature embedding corresponding to each subject. The proposed LG-GNN is thoroughly validated using two public datasets for ASD and AD classification. The experimental results demonstrated that it achieves the state-of-the-art performance in terms of various evaluation metrics.
Hao Zhang 0113, Ran Song 0001, Lin Zhang 0041, Dawei Wang 0015, Cong Wang 0007, Wei Zhang 0021
IEEE Trans. Medical Imaging7
2023 Circular Accessible Depth: A Robust Traversability Representation for UGV Navigation
abstract
In this article, we present the circular accessible depth (CAD), a robust traversability representation for an unmanned ground vehicle (UGV) to learn traversability in various scenarios containing irregular obstacles. To predict CAD, we propose a neural network, namely CADNet, with an attention-based multiframe point cloud fusion module, stability-attention module (SAM), to encode the spatial features from point clouds captured by LiDAR. CAD is designed based on the polar coordinate system and focuses on predicting the border of traversable area. Since it encodes the spatial information of the surrounding environment, which enables a semisupervised learning for the CADNet, and thus, desirably avoids annotating a large amount of data. Extensive experiments demonstrate that CAD outperforms baselines in terms of robustness and precision. We also implement our method on a real UGV and show that it performs well in real-world scenarios.
Shikuan Xie, Ran Song 0001, Yuenan Zhao, Xueqin Huang, Yibin Li 0001, Wei Zhang 0021
IEEE Trans. Robotics6
2023 Watch and Act: Learning Robotic Manipulation From Visual Demonstration
abstract
Learning from demonstration holds the promise of enabling robots to learn diverse actions from expert experience. In contrast to learning from observation-action pairs, humans learn to imitate in a more flexible and efficient manner: learning behaviors by simply “watching.” In this article, we propose a “watch-and-act” imitation learning pipeline that endows a robot with the ability of learning diverse manipulations from visual demonstrations. Specifically, we address this problem by intuitively casting it as two subtasks: 1) understanding the demonstration video and 2) learning the demonstrated manipulations. First, a captioning module based on visual change is presented to understand the demonstration by translating the demonstration video into a command sentence. Then, to execute the captioning command, a manipulation module that learns the demonstrated manipulations is built upon an instance segmentation model and a manipulation affordance prediction model. We validate the superiority of the two modules over existing methods separately via extensive experiments and demonstrate the whole robotic imitation system developed based on the two modules in diverse scenarios using a real robotic arm. Supplementary video is available athttps://vsislab.github.io/watch-and-act/.
Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Hesheng Wang 0001, Yibin Li 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Visual Consensus Modeling for Video-Text Retrieval
abstract
In this paper, we propose a novel method to mine the commonsense knowledge shared between the video and text modalities for video-text retrieval, namely visual consensus modeling. Different from the existing works, which learn the video and text representations and their complicated relationships solely based on the pairwise video-text data, we make the first attempt to model the visual consensus by mining the visual concepts from videos and exploiting their co-occurrence patterns within the video and text modalities with no reliance on any additional concept annotations. Specifically, we build a shareable and learnable graph as the visual consensus, where the nodes denoting the mined visual concepts and the edges connecting the nodes representing the co-occurrence relationships between the visual concepts. Extensive experimental results on the public benchmark datasets demonstrate that our proposed method, with the ability to effectively model the visual consensus, achieves state-of-the-art performances on the bidirectional video-text retrieval task. Our code is available at https://github.com/sqiangcao99/VCM.
Shuqiang Cao, Bairui Wang, Wei Zhang 0021, Lin Ma 0002
AAAI3
2022 Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence Grounding
abstract
Weakly supervised temporal sentence grounding aims to temporally localize the target segment corresponding to a given natural language query, where it provides video-query pairs without temporal annotations during training. Most existing methods use the fused visual-linguistic feature to reconstruct the query, where the least reconstruction error determines the target segment. This work introduces a novel approach that explores the inter-contrast between videos in a composed video by selecting components from two different videos and fusing them into a single video. Such a straightforward yet effective composition strategy provides the temporal annotations at multiple composed positions, resulting in numerous videos with temporal ground-truths for training the temporal sentence grounding task. A transformer framework is introduced with multi-tasks training to learn a compact but efficient visual-linguistic space. The experimental results on the public Charades-STA and ActivityNet-Caption dataset demonstrate the effectiveness of the proposed method, where our approach achieves comparable performance over the state-of-the-art weakly-supervised baselines. The code is available at https://github.com/PPjmchen/Composition_WSTG.
Jiaming Chen 0001, Weixin Luo, Wei Zhang 0021, Lin Ma 0002
AAAI3
2022 Unsupervised Multi-View CNN for Salient View Selection and 3D Interest Point Detection
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu
Int. J. Comput. Vis.2
2022 3D Layout Estimation via Weakly Supervised Learning of Plane Parameters From 2D Segmentation
abstract
The task of 3D layout estimation in an indoor scene is to predict the holistic 3D structural information of the scene from an RGB image. It is costly to obtain the ground truth 3D layout, and this issue severely restricts the learning based 3D layout estimation approaches. In this paper, we present a novel weakly supervised learning framework that is able to learn the 3D layout effectively with 2D layout segmentation mask as supervision. We employ a deep neural network to predict the plane parameters and camera intrinsic parameters in the image. Based on the predicted plane instances, the 3D layout as well as the corresponding depth map and 2D segmentation can be generated. The key objectives for learning meaningful plane parameters are the label consistency of layout segmentation and depth consistency of border pixels from adjacent planes, with which the ground truth 2D layout segmentation is able to supervise the learning of the 3D layout. We further incorporate 3D geometric reasoning and prior knowledge in the learning process to ensure that the learned 3D layout is realistic and reasonable. Experimental results show that our method can produce accurate 3D layout estimates by weakly supervised learning.
Weidong Zhang 0005, Youmei Zhang, Ran Song 0001, Ying Liu 0026, Wei Zhang 0021
IEEE Trans. Image Process.5
2021 UAV-Human: A Large Benchmark for Human Behavior Understanding With Unmanned Aerial Vehicles
abstract
Human behavior understanding with unmanned aerial vehicles (UAVs) is of great significance for a wide range of applications, which simultaneously brings an urgent demand of large, challenging, and comprehensive benchmarks for the development and evaluation of UAV-based models. However, existing benchmarks have limitations in terms of the amount of captured data, types of data modalities, categories of provided tasks, and diversities of subjects and environments. Here we propose a new benchmark - UAV-Human - for human behavior understanding with UAVs, which contains 67,428 multi-modal video sequences and 119 subjects for action recognition, 22,476 frames for pose estimation, 41,290 frames and 1,144 identities for person re-identification, and 22,263 frames for attribute recognition. Our dataset was collected by a flying UAV in multiple urban and rural districts in both daytime and night-time over three months, hence covering extensive diversities w.r.t subjects, backgrounds, illuminations, weathers, occlusions, camera motions, and UAV flying attitudes. Such a comprehensive and challenging benchmark shall be able to promote the research of UAV-based human behavior understanding, including action recognition, pose estimation, re-identification, and attribute recognition. Furthermore, we propose a fisheye-based action recognition method that mitigates the distortions in fisheye videos via learning unbounded transformations guided by flat RGB videos. Experiments show the efficacy of our method on the UAV-Human dataset.
Jun Liu 0036, Wei Zhang 0021, Yun Ni, Zhiheng Li 0005
CVPR3
2021 Mesh Saliency: An Independent Perceptual Measure or a Derivative of Image Saliency?
abstract
While mesh saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and is well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art mesh saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that mesh saliency might associate with the saliency of 2D natural images. This paper proposes a novel deep neural network for learning mesh saliency using image saliency ground truth to 1) investigate whether mesh saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting mesh saliency. Through extensive experiments, we not only demonstrate that our method outperforms the current state-of-the-art mesh saliency method by 116% and 21% in terms of linear correlation coefficient and AUC respectively, but also reveal that mesh saliency is intrinsically related with both image saliency and object categorical information. Codes are available at https://github.com/rsong/MIMO-GAN.
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin
CVPR2
2021 Autonomous Multi-View Navigation via Deep Reinforcement Learning
abstract
In this paper, we propose a novel deep reinforcement learning (DRL) system for the autonomous navigation of mobile robots that consists of three modules: map navigation, multi-view perception and multi-branch control. Our DRL system takes as the input a routed map provided by a global planner and three RGB images captured by a multi-camera setup to gather global and local information, respectively. In particular, we present a multi-view perception module based on an attention mechanism to filter out redundant information caused by multi-camera sensing. We also replace raw RGB images with low-dimensional representations via a specifically designed network, which benefits a more robust sim2real transfer learning. Extensive experiments in both simulated and real-world scenarios demonstrate that our system outperforms state-of-the-art approaches.
Xueqin Huang, Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Yibin Li 0001
ICRA3
2021 A Hierarchical Framework for Quadruped Locomotion Based on Reinforcement Learning
abstract
Quadruped locomotion is a challenging task for learning-based algorithms. It requires tedious manual tuning and is difficult to deploy in reality due to the reality gap. In this paper, we propose a quadruped robot learning system for agile locomotion which does not require any pre-training and works well in various real-world terrains. We introduce a hierarchical learning framework that uses reinforcement learning as the high-level policy to adjust the low-level trajectory generator for better adaptability to the terrain. We compact the observation and action space of the reinforcement learning to deploy it on a host computer in reality. Besides, we design a trajectory generator guided by robot posture, which can generate adaptive foot trajectory to interact with the environment. Experimental results show that our system can be easily deployed in reality while only trained in simulation, and also has the advantages of fast convergence and good terrain adaptability. The supplementary video demonstration is available at https://vsislab.github.io/hfql/.
Wenhao Tan, Wei Zhang 0021, Ran Song 0001, Yu Zheng 0001, Yibin Li 0001
IROS3
2021 PackerBot: Variable-Sized Product Packing with Heuristic Deep Reinforcement Learning
abstract
Product packing is a typical application in ware-house automation that aims to pick objects from unstructured piles and place them into bins with optimized placing policy. However, it still remains a significant challenge to finish the product packing tasks in general logistics scenarios where the objects are variable-sized and the configurations are complex. In this work, we present the PackerBot, a complete robotic pipeline for performing variable-sized product packing in unstructured scenes. First, by leveraging the imperfect experience of human packer, we propose a heuristic DRL framework for learning optimal online 3D bin packing policy. Then we integrate it with a 6-DoF suction-based picking module and a product size estimation module, leading to a complete product packing system, namely the PackerBot. Extensive experimental results show that our method achieves the state-of-the-art performance in both simulated and real-world tests. The video demonstration is available at: https://vsislab.github.io/packerbot.
Zifei Yang, Shuai Song, Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Yibin Li 0001
IROS4
2021 PoT-GAN: Pose Transform GAN for Person Image Synthesis
abstract
Pose-based person image synthesis aims to generate a new image containing a person with a target pose conditioned on a source image containing a person with a specified pose. It is challenging as the target pose is arbitrary and often significantly differs from the specified source pose, which leads to large appearance discrepancy between the source and the target images. This paper presents the Pose Transform Generative Adversarial Network (PoT-GAN) for person image synthesis where the generator explicitly learns the transform between the two poses by manipulating the corresponding multi-scale feature maps. By incorporating the learned pose transform information into the multi-scale feature maps of the source image in a GAN architecture, our method reliably transfers the appearance of the person in the source image to the target pose with no need for any hard-coded spatial information depicting the change of pose. According to both qualitative and quantitative results, the proposed PoT-GAN demonstrates a state-of-the-art performance on three publicly available datasets for person image synthesis.
Wei Zhang 0021, Ran Song 0001, Zhiheng Li 0005, Jun Liu 0036, Xiaolei Li 0003, Shijian Lu
IEEE Trans. Image Process.2
2021 BATCH: A Scalable Asymmetric Discrete Cross-Modal Hashing
abstract
Supervised cross-modal hashing has attracted much attention. However, there are still some challenges, e.g., how to effectively embed the label information into binary codes, how to avoid using a large similarity matrix and make a model scalable to large-scale datasets, how to efficiently solve the binary optimization problem. To address these challenges, in this paper, we present a novel supervised cross-modal hashing method, i.e., scalaBle Asymmetric discreTe Cross-modal Hashing, BATCH for short. It leverages collective matrix factorization to learn a common latent space for the labels and different modalities, and embeds the labels into binary codes by minimizing a distance-distance difference problem. Furthermore, it builds a connection between the common latent space and the hash codes by an asymmetric strategy. In the light of this, it can perform cross-modal retrieval and embed more similarity information into the binary codes. In addition, it introduces a quantization minimization term and orthogonal constraints into the optimization problem, and generates the binary codes discretely. Therefore, the quantization error and redundancy may be much reduced. Moreover, it is a two-step method, making the optimization simple and scalable to large-scale datasets. Extensive experimental results on three benchmark datasets demonstrate that BATCH outperforms some state-of-the-art cross-modal hashing methods in terms of accuracy and efficiency.
Yongxin Wang 0001, Xin Luo 0006, Liqiang Nie, Jingkuan Song, Wei Zhang 0021, Xin-Shun Xu
IEEE Trans. Knowl. Data Eng.5
2021 From Edge to Keypoint: An End-to-End Framework For Indoor Layout Estimation
abstract
The task of spatial layout estimation of monocular image is to segment an RGB image of indoor scenes with semantic surface labels (i.e., ceiling, floor, front wall, left wall, and right wall). Most recent methods have to produce layout hypotheses based on the estimated edge map or semantic labels, and then rank the layout hypotheses. In this paper, we present an end-to-end framework that can directly output the layout type and keypoint coordinates (defined in the LSUN challenge). The proposed method takes advantage of transfer learning via learning on the fake samples, i.e., plenty of artificial {type, keypoints, edge map} triplets are generated to learn the mapping from edge maps to keypoint coordinates. Generative adversarial network (GAN) is implemented in this work for domain adaptation of the edge maps. Experimental results show that the proposed method can achieve state-of-the-art layout estimation performance on benchmark datasets.
Weidong Zhang 0005, Qian Zhang 0076, Wei Zhang 0021, Jason Gu, Yibin Li 0001
IEEE Trans. Multim.3
2021 Visual Navigation With Multiple Goals Based on Deep Reinforcement Learning
abstract
Learning to adapt to a series of different goals in visual navigation is challenging. In this work, we present a model-embedded actor-critic architecture for the multigoal visual navigation task. To enhance the task cooperation in multigoal learning, we introduce two new designs to the reinforcement learning scheme: inverse dynamics model (InvDM) and multigoal colearning (MgCl). Specifically, InvDM is proposed to capture the navigation-relevant association between state and goal and provide additional training signals to relieve the sparse reward issue. MgCl aims at improving the sample efficiency and supports the agent to learn from unintentional positive experiences. Besides, to further improve the scene generalization capability of the agent, we present an enhanced navigation model that consists of two self-supervised auxiliary task modules. The first module, which is named path closed-loop detection, helps to understand whether the state has been experienced. The second one, namely the state-target matching module, tries to figure out the difference between state and goal. Extensive results on the interactive platform AI2-THOR demonstrate that the agent trained with the proposed method converges faster than state-of-the-art methods while owning good generalization capability. The video demonstration is available at https://vsislab.github.io/mgvn.
Zhenhuan Rao, Yuechen Wu, Zifei Yang, Wei Zhang 0021, Shijian Lu, Weizhi Lu, Zhengjun Zha
IEEE Trans. Neural Networks Learn. Syst.4
2021 Spatial-temporal Regularized Multi-modality Correlation Filters for Tracking with Re-detection
abstract
The development of multi-spectrum image sensing technology has brought great interest in exploiting the information of multiple modalities (e.g., RGB and infrared modalities) for solving computer vision problems. In this article, we investigate how to exploit information from RGB and infrared modalities to address two important issues in visual tracking: robustness and object re-detection. Although various algorithms that attempt to exploit multi-modality information in appearance modeling have been developed, they still face challenges that mainly come from the following aspects: (1) the lack of robustness to deal with large appearance changes and dynamic background, (2) failure in re-capturing the object when tracking loss happens, and (3) difficulty in determining the reliability of different modalities. To address these issues and perform effective integration of multiple modalities, we propose a new tracking-by-detection algorithm called Adaptive Spatial-temporal Regulated Multi-Modality Correlation Filter. Particularly, an adaptive spatial-temporal regularization is imposed into the correlation filter framework in which the spatial regularization can help to suppress effect from the cluttered background while the temporal regularization enables the adaptive incorporation of historical appearance cues to deal with appearance changes. In addition, a dynamic modality weight learning algorithm is integrated into the correlation filter training, which ensures that more reliable modalities gain more importance in target tracking. Experimental results demonstrate the effectiveness of the proposed method.
Xiangyuan Lan, Zifei Yang, Wei Zhang 0021, Pong C. Yuen
ACM Trans. Multim. Comput. Commun. Appl.3
2020 GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes
Weidong Zhang 0005, Wei Zhang 0021, Yinda Zhang 0001
ECCV (16)2
2020 HARD-Net: Hardness-AwaRe Discrimination Network for 3D Early Activity Prediction
Jun Liu 0036, Wei Zhang 0021, Ling-Yu Duan
ECCV (11)3
2020 Unsupervised Multi-view CNN for Salient View Selection of 3D Objects and Scenes
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu
ECCV (19)2
2020 Cross-context Visual Imitation Learning from Demonstrations
abstract
Imitation learning enables robots to learn a task by simply watching the demonstration of the task. Current imitation learning methods usually require the learner and demonstrator to occur in the same context. This limits their scalability to practical applications. In this paper, we propose a more general imitation learning method which allows the learner and the demonstrator to come from different contexts, such as different viewpoints, backgrounds, and object positions and appearances. Specifically, we design a robotic system consisting of three models: context translation model, depth prediction model and multi-modal inverse dynamics model. First, the context translation model translates the demonstration to the context of learner from a different context. Then combining the color observation and depth observation as inputs, the inverse model maps the multi-modal observations into actions to reproduce the demonstration, where the depth observation is provided by a depth prediction model. By performing the block stacking tasks both in simulation and real world, we prove the cross-context learning advantage of the proposed robotic system over other systems.
Wei Zhang 0021, Weizhi Lu, Hesheng Wang 0001, Yibin Li 0001
ICRA2
2020 Grasp for Stacking via Deep Reinforcement Learning
abstract
Integrated robotic arm system should contain both grasp and place actions. However, most grasping methods focus more on how to grasp objects, while ignoring the placement of the grasped objects, which limits their applications in various industrial environments. In this research, we propose a model-free deep Q-learning method to learn the grasping-stacking strategy end-to-end from scratch. Our method maps the images to the actions of the robotic arm through two deep networks: the grasping network (GNet) using the observation of the desk and the pile to infer the gripper's position and orientation for grasping, and the stacking network (SNet) using the observation of the platform to infer the optimal location when placing the grasped object. To make a long-range planning, the two observations are integrated in the grasping for stacking network (GSN). We evaluate the proposed GSN on a grasping-stacking task in both simulated and real-world scenarios.
Wei Zhang 0021, Ran Song 0001, Lin Ma 0002, Yibin Li 0001
ICRA2
2020 Learn by Observation: Imitation Learning for Drone Patrolling from Videos of A Human Navigator
abstract
We present an imitation learning method for autonomous drone patrolling based only on raw videos. Different from previous methods, we propose to let the drone learn patrolling in the air by observing and imitating how a human navigator does it on the ground. The observation process enables the automatic collection and annotation of data using inter-frame geometric consistency, resulting in less manual effort and high accuracy. Then a newly designed neural network is trained based on the annotated data to predict appropriate directions and translations for the drone to patrol in a lane-keeping manner as humans. Our method allows the drone to fly at a high altitude with a broad view and low risk. It can also detect all accessible directions at crossroads and further carry out the integration of available user instructions and autonomous patrolling control commands. Extensive experiments are conducted to demonstrate the accuracy of the proposed imitating learning process as well as the reliability of the holistic system for autonomous drone navigation. The codes, datasets as well as video demonstrations are available at https://vsislab.github.io/uavpatrol.
Shilei Chu, Wei Zhang 0021, Ran Song 0001, Yibin Li 0001
IROS3
2020 Autonomous Robot Navigation Based on Multi-Camera Perception
abstract
In this paper, we propose an autonomous method for robot navigation based on a multi-camera setup that takes advantage of a wide field of view. A new multi-task network is designed for handling the visual information supplied by the left, central and right cameras to find the passable area, detect the intersection and infer the steering. Based on the outputs of the network, three navigation indicators are generated and then combined with the high-level control commands extracted by the proposed MapNet, which are finally fed into the driving controller. The indicators are also used through the controller for adjusting the driving velocity, which assists the robot to adjust the speed for smoothly bypassing obstacles. Experiments in real-world environments demonstrate that our method performs well in both local obstacle avoidance and global goal-directed navigation tasks.
Kunyan Zhu, Wei Zhang 0021, Ran Song 0001, Yibin Li 0001
IROS3
2020 Online Decision Based Visual Tracking via Reinforcement Learning
abstract
A deep visual tracker is typically based on either object detection or template matching while each of them is only suitable for a particular group of scenes. It is straightforward to consider fusing them together to pursue more reliable tracking. However, this is not wise as they follow different tracking principles. Unlike previous fusion-based methods, we propose a novel ensemble framework, named DTNet, with an online decision mechanism for visual tracking based on hierarchical reinforcement learning. The decision mechanism substantiates an intelligent switching strategy where the detection and the template trackers have to compete with each other to conduct tracking within different scenes that they are adept in. Besides, we present a novel detection tracker which avoids the common issue of incorrect proposal. Extensive results show that our DTNet achieves state-of-the-art tracking performance as well as good balance between accuracy and efficiency. The project website is available at https://vsislab.github.io/DTNet/.
Ke Song 0003, Wei Zhang 0021, Ran Song 0001, Yibin Li 0001
NeurIPS2
2020 Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
abstract
In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a reconstruction network (RecNet) in a novel encoder-decoder-reconstructor architecture, which leverages both forward (video to sentence) and backward (sentence to video) flows for video captioning. Specifically, the encoder-decoder component makes use of the forward flow to produce a sentence based on the encoded video semantic features. Two types of reconstructors are subsequently proposed to employ the backward flow and reproduce the video features from local and global perspectives, respectively, capitalizing on the hidden state sequence generated by the decoder. Moreover, in order to make a comprehensive reconstruction of the video features, we propose to fuse the two types of reconstructors together. The generation loss yielded by the encoder-decoder component and the reconstruction loss introduced by the reconstructor are jointly cast into training the proposed RecNet in an end-to-end fashion. Furthermore, the RecNet is fine-tuned by CIDEr optimization via reinforcement learning, which significantly boosts the captioning performance. Experimental results on benchmark datasets demonstrate that the proposed reconstructor can boost the performance of video captioning consistently.
Wei Zhang 0021, Bairui Wang, Lin Ma 0002, Wei Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 SCRATCH: A Scalable Discrete Matrix Factorization Hashing Framework for Cross-Modal Retrieval
abstract
In this paper, we present a novel supervised cross-modal hashing framework, namely Scalable disCRete mATrix faCtorization Hashing (SCRATCH). First, it utilizes collective matrix factorization on original features together with label semantic embedding, to learn the latent representations in a shared latent space. Thereafter, it generates binary hash codes based on the latent representations. During optimization, it avoids using a large n × n similarity matrix and generates hash codes discretely. Besides, based on different objective functions, learning strategy, and features, we further present three models in this framework, i.e., SCRATCH-o, SCRATCH-t, and SCRATCH-d. The first one is a one-step method, learning the hash functions and the binary codes in the same optimization problem. The second is a two-step method, which first generates the binary codes and then learns the hash functions based on the learned hash codes. The third one is a deep version of SCRATCH-t, which utilizes deep neural networks as hash functions. The extensive experiments on two widely used benchmark datasets demonstrate that SCRATCH-o and SCRATCH-t outperform some state-of-the-art shallow hashing methods for cross-modal retrieval. The SCRATCH-d also outperforms some state-of-the-art deep hashing models.
Zhen-Duo Chen 0001, Chuan-Xiang Li, Xin Luo 0006, Liqiang Nie, Wei Zhang 0021, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.5
2020 Visual Object Tracking via Guessing and Matching
abstract
Visual object tracking is a fundamental and time-critical vision task. However, most trackers such as SiamFC and CFNet missed the object movement and simply defined the searching region centered at the location of the target in the previous frame. So they tend to fail in the cases with severe occlusion or a large displacement of the target. In this paper, we consider the object tracking as a dual-task problem of guessing and matching. A guess module is to estimate the motion trend of the target by reinforcement learning based on the observations on appearance changes and motion history. Rather than using the previous location of the target, we may have a more accurate center to locate the searching region. Benefited from such improved searching region, the match module becomes less prone to the object drift problem, and can easily identify the target from the potential distractors in the background. Extensive experimental results on benchmark datasets such as RGBT, OTB-2013, OTB-50 and OTB-100, show that the proposed method achieves leading performance compared to state-of-the-art trackers. Moreover, the proposed tracker could maintain real-time speed, giving itself the potential in practical applications.
Ke Song 0003, Wei Zhang 0021, Weizhi Lu, Zhengjun Zha, Xiangyang Ji, Yibin Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Edge-Semantic Learning Strategy for Layout Estimation in Indoor Environment
abstract
Visual cognition of the indoor environment can benefit from the spatial layout estimation, which is to represent an indoor scene with a 2-D box on a monocular image. In this paper, we propose to fully exploit the edge and semantic information of a room image for layout estimation. More specifically, we present an encoder-decoder network with shared encoder and two separate decoders, which are composed of multiple deconvolution (transposed convolution) layers, to jointly learn the edge maps and semantic labels of a room image. We combine these two network predictions in a scoring function to evaluate the quality of the layouts, which are generated by ray sampling and from a predefined layout pool. Guided by the scoring function, we apply a novel refinement strategy to further optimize the layout hypotheses. Experimental results show that the proposed network can yield accurate estimates of edge maps and semantic labels. By fully utilizing the two different types of labels, the proposed method achieves the state-of-the-art layout estimation performance on the benchmark datasets.
Weidong Zhang 0005, Wei Zhang 0021, Jason Gu
IEEE Trans. Cybern.2
2020 A Multi-Scale Spatial-Temporal Attention Model for Person Re-Identification in Videos
abstract
In this paper, we propose a novel deep neural network based attention model to learn the representative local regions from a video sequence for person re-identification. Specifically, we propose a multi-scale spatial-temporal attention (MSTA) model to measure the regions of each frame in different scales from the perspective of whole video sequence. Compared to traditional temporal attention models, MSTA focuses on exploiting the importance of local regions of each frame to the whole video representation in both spatial and temporal domains. A new training strategy is designed for the proposed model by incorporating the image-to-image mode with the videoto- video mode. Extensive experiments on benchmark datasets demonstrate the superiority of the proposed model over state-ofthe- art methods.
Wei Zhang 0021, Xuanyu He, Weizhi Lu, Zhengjun Zha, Qi Tian 0001
IEEE Trans. Image Process.1
2020 Exploring Discriminative Representations for Image Emotion Recognition With CNNs
abstract
Image emotion recognition aims to automatically categorize the emotion conveyed by an image. The potential of deep representation has been demonstrated in recent research on image emotion recognition. To better understand how CNNs work in emotion recognition, we investigate the deep features by visualizing them in this work. This study shows that the deep models mainly rely on the image content but miss the image style information such as color, texture, and shapes that are low-level visual features but are vital for evoking emotions. To form a more discriminative representation for emotion recognition, we propose a novel CNN model that learns and integrates the content information from the high layers of the deep network with the style information from the lower layers. The uncertainty of image emotion labels is also investigated in this paper. Rather than using the emotion labels for training directly, as in previous work, a new loss function is designed by including the emotion labeling quality to optimize the proposed inference model. Extensive experiments on benchmark datasets are conducted to demonstrate the superiority of the proposed representation.
Wei Zhang 0021, Xuanyu He, Weizhi Lu
IEEE Trans. Multim.1
2019 Hierarchical Photo-Scene Encoder for Album Storytelling
abstract
In this paper, we propose a novel model with a hierarchical photo-scene encoder and a reconstructor for the task of album storytelling. The photo-scene encoder contains two subencoders, namely the photo and scene encoders, which are stacked together and behave hierarchically to fully exploit the structure information of the photos within an album. Specifically, the photo encoder generates semantic representation for each photo while exploiting temporal relationships among them. The scene encoder, relying on the obtained photo representations, is responsible for detecting the scene changes and generating scene representations. Subsequently, the decoder dynamically and attentively summarizes the encoded photo and scene representations to generate a sequence of album representations, based on which a story consisting of multiple coherent sentences is generated. In order to fully extract the useful semantic information from an album, a reconstructor is employed to reproduce the summarized album representations based on the hidden states of the decoder. The proposed model can be trained in an end-to-end manner, which results in an improved performance over the state-of-the-arts on the public visual storytelling (VIST) dataset. Ablation studies further demonstrate the effectiveness of the proposed hierarchical photo-scene encoder and reconstructor.
Bairui Wang, Lin Ma 0002, Wei Zhang 0021
AAAI3
2019 Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion Network
abstract
In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We construct a novel gated fusion network, with one particularly designed cross-gating (CG) block, to effectively encode and fuse different types of representations, e.g., the motion and content features of an input video. One POS sequence generator relies on this fused representation to predict the global syntactic structure, which is thereafter leveraged to guide the video captioning generation and control the syntax of the generated sentence. Specifically, a gating strategy is proposed to dynamically and adaptively incorporate the global syntactic POS information into the decoder for generating each word. Experimental results on two benchmark datasets, namely MSR-VTT and MSVD, demonstrate that the proposed model can well exploit complementary information from multiple representations, resulting in improved performances. Moreover, the generated global POS information can well capture the global syntactic structure of the sentence, and thus be exploited to control the syntactic structure of the description. Such POS information not only boosts the video captioning performance but also improves the diversity of the generated captions. Our code is at: https://github.com/vsislab/Controllable_XGating.
Bairui Wang, Lin Ma 0002, Wei Zhang 0021, Jingwen Wang 0003, Wei Liu 0005
ICCV3
2019 Road Context-Aware Intrusion Detection System for Autonomous Cars
Jingxuan Jiang, Chundong Wang 0001, Sudipta Chattopadhyay 0001, Wei Zhang 0021
ICICS4
2019 Attention to Head Locations for Crowd Counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot, Wei Zhang 0021
ICIG (2)5
2019 Exploring the Task Cooperation in Multi-goal Visual Navigation
abstract
Learning to adapt to a series of different goals in visual navigation is challenging. In this work, we present a model-embedded actor-critic architecture for the multi-goal visual navigation task. To enhance the task cooperation in multi-goal learning, we introduce two new designs to the reinforcement learning scheme: inverse dynamics model (InvDM) and multi-goal co-learning (MgCl). Specifically, InvDM is proposed to capture the navigation-relevant association between state and goal, and provide additional training signals to relieve the sparse reward issue. MgCl aims at improving the sample efficiency and supports the agent to learn from unintentional positive experiences. Extensive results on the interactive platform AI2-THOR demonstrate that the proposed method converges faster than state-of-the-art methods while producing more direct routes to navigate to the goal. The video demonstration is available at: https://youtube.com/channel/UCtpTMOsctt3yPzXqe_JMD3w/videos.
Yuechen Wu, Zhenhuan Rao, Wei Zhang 0021, Shijian Lu, Weizhi Lu, Zhengjun Zha
IJCAI3
2019 MSR: Multi-Scale Shape Regression for Scene Text Detection
abstract
State-of-the-art scene text detection techniques predict quadrilateral boxes that are prone to localization errors while dealing with straight or curved text lines of different orientations and lengths in scenes. This paper presents a novel multi-scale shape regression network (MSR) that is capable of locating text lines of different lengths, shapes and curvatures in scenes. The proposed MSR detects scene texts by predicting dense text boundary points that inherently capture the location and shape of text lines accurately and are also more tolerant to the variation of text line length as compared with the state of the arts using proposals or segmentation. Additionally, the multi-scale network extracts and fuses features at different scales which demonstrates superb tolerance to the text scale variation. Extensive experiments over several public datasets show that the proposed MSR obtains superior detection performance for both curved and straight text lines of different lengths and orientations.
Chuhui Xue, Shijian Lu, Wei Zhang 0021
IJCAI3
2019 Learning Actions from Human Demonstration Video for Robotic Manipulation
abstract
Learning actions from human demonstration is an emerging trend for designing intelligent robotic systems, which can be referred as video to command. The performance of such approach highly relies on the quality of video captioning. However, the general video captioning methods focus more on the understanding of the full frame, lacking of consideration on the specific object of interests in robotic manipulations. We propose a novel deep model to learn actions from human demonstration video for robotic manipulation. It consists of two deep networks, grasp detection network (GNet) and video captioning network (CNet). GNet performs two functions: providing grasp solutions and extracting the local features for the object of interests in robotic manipulation. CNet outputs the captioning results by fusing the features of both full frames and local objects. Experimental results on UR5 robotic arm show that our method could produce more accurate command from video demonstration than state-of-the-art work, thereby leading to more robust grasping performance.
Wei Zhang 0021, Weizhi Lu, Hesheng Wang 0001, Yibin Li 0001
IROS2
2019 Progressive Retinex: Mutually Reinforced Illumination-Noise Perception Network for Low-Light Image Enhancement
abstract
Contrast enhancement and noise removal are coupled problems for low-light image enhancement. The existing Retinex based methods do not take the coupling relation into consideration, resulting in under or over-smoothing of the enhanced images. To address this issue, this paper presents a novel progressive Retinex framework, in which illumination and noise of low-light image are perceived in a mutually reinforced manner, leading to noise reduction low-light enhancement results. Specifically, two fully pointwise convolutional neural networks are devised to model the statistical regularities of ambient light and image noise respectively, and to leverage them as constraints to facilitate the mutual learning process. The proposed method not only suppresses the interference caused by the ambiguity between tiny textures and image noises, but also greatly improves the computational efficiency. Moreover, to solve the problem of insufficient training data, we propose an image synthesis strategy based on camera imaging model, which generates color images corrupted by illumination-dependent noises. Experimental results on both synthetic and real low-light images demonstrate the superiority of our proposed approaches against the State-Of-The-Art (SOTA) low-light enhancement methods.
Yang Wang 0015, Yang Cao 0010, Zhengjun Zha, Jing Zhang 0037, Zhiwei Xiong, Wei Zhang 0021, Feng Wu 0001
ACM Multimedia6
2019 Illumination-Invariant Person Re-Identification
abstract
Due to the effect of weak illumination, person images captured by surveillance cameras usually contain various degradations such as color shift, low contrast and noise. These degradations result in severe discriminant information loss, which makes the person re-identification (re-id) more challenging. However, existing person re-identification approaches are designed based on the assumption that the pedestrians images are under well lighting conditions, which is impractical in real-world scenarios. Inspired by the Retinex theory, we propose a illumination-invariant person re-identification framework which is able to simultaneously achieve Retinex illumination decomposition and person re-identification. We first verify that directly using weak illuminated images can greatly reduce the performance of person re-id. We then design a bottom-up attention network to remove the effect of weak illumination and obtain the enhanced image without introducing over-enhancement. To effectively connect low-level and high-level vision tasks, a joint training strategy is further introduced to boost the performance of person re-id under weak illumination conditions. Experiments have demonstrated the advantages of our method on benchmarks with severe lighting changes and low light conditions.
Zhengjun Zha, Xueyang Fu, Wei Zhang 0021
ACM Multimedia4
2019 Distributed Caching Popular Services by Using Deep Q-Learning in Converged Networks
abstract
Content caching offers an effective solution to reduce the traffic load and alleviate the burden on backhaul links in future wireless networks. In this paper, we study the converged networks to push and cache the popular services. The popular services are delivered by the broadcasting networks, and cached in the router nodes in a distributed cache network. Due to the limited storage capacity of the router node, we formulate the service scheduling problem as a Markov Decision Process (MDP), aiming to maximize the equivalent throughput. Considering the large state space involved in the distributed cache network, it is great challenge to obtain a tractable solution by the classical optimization algorithm. To tackle this problem, we propose a service scheduling strategy based on deep Q-learning. Simulation results demonstrate that the proposed scheme can significantly improve the equivalent throughput of the converged networks.
Yuzhe Fang, Jian Xiong 0001, Peng Cheng 0002, Wei Zhang 0021
VTC Fall4
2019 Salient object detection with adversarial training
abstract
The generative adversarial network has been shown to produce state‐of‐the‐art results of image generation. In this study, the authors propose a novel adversarial training method to train salient object detection (SOD) models. They train a convolutional SOD network along with a gated adversarial network that discriminates salient maps coming either from the ground truth or from the SOD network. The motivation for our approach is that the adversarial network can detect and correct pixel‐wise errors between ground truth salient detection maps and the ones produced by the convolutional network. Our experiments show that the adversarial training approach leads to state‐of‐the‐art performance on MSRA‐B, extended complex scene saliency dataset, HKU‐IS, DUT, and SOD dataset.
Zhijie Wang 0010, Wei Zhang 0021, Xuewen Rong, Yibin Li 0001
IET Image Process.2
2019 Special issue on deep learning for intelligent sensing, decision-making and control
Wei Zhang 0021, Junchi Yan, Zhiyong Liu 0001, Zhigang Zeng
Neurocomputing1
2019 Reversible data hiding for high dynamic range images using edge information
Xuanyu He, Wei Zhang 0021, Lin Ma 0002, Yibin Li 0001
Multim. Tools Appl.2
2019 Real-time visual tracking with ELM augmented adaptive correlation filter
Kuixiang Liu, Baochen Yao, Jun Tang 0007, Wei Zhang 0021
Pattern Recognit. Lett.5
2019 Coarse-to-Fine UAV Target Tracking With Deep Reinforcement Learning
abstract
The aspect ratio of a target changes frequently during an unmanned aerial vehicle (UAV) tracking task, which makes the aerial tracking very challenging. Traditional trackers struggle from such a problem as they mainly focus on the scale variation issue by maintaining a certain aspect ratio. In this paper, we propose a coarse-to-fine deep scheme to address the aspect ratio variation in UAV tracking. The coarse-tracker first produces an initial estimate for the target object, then a sequence of actions are learned to fine-tune the four boundaries of the bounding box. The coarse-tracker and the fine-tracker are designed to have different action spaces and operating target. The former dominates the entire bounding box and the latter focuses on the refinement of each boundary. They are trained jointly by sharing the perception network with an end-to-end reinforcement learning architecture. Experimental results on benchmark aerial data set prove that the proposed approach outperforms existing trackers and produces significant accuracy gains in dealing with the aspect ratio variation in UAV tracking.
Wei Zhang 0021, Ke Song 0003, Xuewen Rong, Yibin Li 0001
IEEE Trans Autom. Sci. Eng.1
2019 Learning Compact Appearance Representation for Video-Based Person Re-Identification
abstract
This paper presents a novel approach for video-based person re-identification using multiple convolutional neural networks (CNNs). Unlike the previous work, we intend to extract a compact yet discriminative appearance representation from several frames rather than the whole sequence. Specifically, given a video, the representative frames are selected based on the walking profile of consecutive frames. A multiple CNN architecture incorporated with feature pooling is proposed to learn and compile the features of the selected representative frames into a compact description about the pedestrian for identification. Experiments are conducted on benchmark data sets to demonstrate the superiority of the proposed method over existing person re-identification approaches.
Wei Zhang 0021, Shengnan Hu, Kan Liu 0001, Zhengjun Zha
IEEE Trans. Circuits Syst. Video Technol.1
2019 Learning Intra-Video Difference for Person Re-Identification
abstract
Siamese networks are prevalent in person re-identification (re-id) tasks to address the similarity and dissimilarity among video frames. It mainly focuses on the inter-video variation between spatio-temporal features extracted from different videos, while the variation between features of the same video has been rarely discussed. In this paper, we introduce the concept of “mean-body” and define an intra-video loss to address the variation between spatio-temporal features of the same video. A novel loss is presented to boost the training of the re-id networks by combining the proposed intra-video loss and the Siamese loss. Specifically, the intra-video loss uses the unique mean-body of each camera viewpoint to make the video sequence more clustered, while the Siamese loss is to make the wrong matching videos more separated. To train the whole network, we update the network and the mean-body in an iterative manner. As a result, the proposed loss is expected to improve the generalization capability of the re-id networks on the testing set. Extensive results demonstrate that the presented approach outperforms the state-of-the-art algorithms on the publicly available data sets, such as PRID2011, iLIDS-VID, and MARS, in terms of re-id accuracy.
Wei Zhang 0021, Weizhi Lu, Xin-Shun Xu, Xiangyang Ji
IEEE Trans. Circuits Syst. Video Technol.1
2019 DECAL: Decomposition-Based Coevolutionary Algorithm for Many-Objective Optimization
abstract
This paper develops a decomposition-based coevolutionary algorithm for many-objective optimization, which evolves a number of subpopulations in parallel for approaching the set of Pareto optimal solutions. The many-objective problem is decomposed into a number of subproblems using a set of well-distributed weight vectors. Accordingly, each subpopulation of the algorithm is associated with a weight vector and is responsible for solving the corresponding subproblem. The exploration ability of the algorithm is improved by using a mating pool that collects elite individuals from the cooperative subpopulations for breeding the offspring. In the subsequent environmental selection, the top-ranked individuals in each subpopulation, which are appraised by aggregation functions, survive for the next iteration. Two new aggregation functions with distinct characteristics are designed in this paper to enhance the population diversity and accelerate the convergence speed. The proposed algorithm is compared with several state-of-the-art many-objective evolutionary algorithms on a large number of benchmark instances, as well as on a real-world design problem. Experimental results show that the proposed algorithm is very competitive.
Yuhui Zhang 0004, Yue-Jiao Gong, Tianlong Gu, Huaqiang Yuan, Wei Zhang 0021, Sam Kwong, Jun Zhang 0003
IEEE Trans. Cybern.5
2019 CAD-Net: A Context-Aware Detection Network for Objects in Remote Sensing Imagery
abstract
Accurate and robust detection of multi-class objects in optical remote sensing images is essential to many real-world applications, such as urban planning, traffic control, searching, and rescuing. However, the state-of-the-art object detection techniques designed for images captured using ground-level sensors usually experience a sharp performance drop when directly applied to remote sensing images, largely due to the object appearance differences in remote sensing images in terms of sparse texture, low contrast, arbitrary orientations, and large-scale variations. This paper presents a novel object detection network [(context-aware detection network (CAD-Net)] that exploits attention-modulated features as well as global and local contexts to address the new challenges in detecting objects from remote sensing images. The proposed CAD-Net learns global and local contexts of objects by capturing their correlations with the global scene (at scene level) and the local neighboring objects or features (at object level), respectively. In addition, it designs a spatial-and-scale-aware attention module that guides the network to focus on more informative regions and features as well as more appropriate feature scales. Experiments over two publicly available object detection data sets for remote sensing images demonstrate that the proposed CAD-Net achieves superior detection performance. The implementation codes will be made publicly available for facilitating future works.
Gongjie Zhang, Shijian Lu, Wei Zhang 0021
IEEE Trans. Geosci. Remote. Sens.3
2019 Feature Aggregation With Reinforcement Learning for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-id) matches two tracks of persons from different cameras. Features are extracted from the images of a sequence and then aggregated as a track feature. Compared to existing works that aggregate frame features by simply averaging them or using temporal models such as recurrent neural networks, we propose an intelligent feature aggregate method based on reinforcement learning. Specifically, we train an agent to determine which frames in the sequence should be abandoned in the aggregation, which can be treated as a decision making process. By this way, the proposed method avoids introducing noisy information of the sequence and retains these valuable frames when generating a track feature. On benchmark data sets, experimental results show that our method can boost the re-id accuracy obviously based on the state-of-the-art models.
Wei Zhang 0021, Xuanyu He, Weizhi Lu, Hong Qiao, Yibin Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Reconstruction Network for Video Captioning
abstract
In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a reconstruction network (RecNet) with a novel encoder-decoder-reconstructor architecture, which leverages both the forward (video to sentence) and backward (sentence to video) flows for video captioning. Specifically, the encoder-decoder makes use of the forward flow to produce the sentence description based on the encoded video semantic features. Two types of reconstructors are customized to employ the backward flow and reproduce the video features based on the hidden state sequence generated by the decoder. The generation loss yielded by the encoder-decoder and the reconstruction loss introduced by the reconstructor are jointly drawn into training the proposed RecNet in an end-to-end fashion. Experimental results on benchmark datasets demonstrate that the proposed reconstructor can boost the encoder-decoder models and leads to significant gains in video caption accuracy.
Bairui Wang, Lin Ma 0002, Wei Zhang 0021, Wei Liu 0005
CVPR3
2018 Diversified Dual Domain-Adversarial Neural Networks
abstract
The application cost machine learning methods often rely on the availability of large-scale data collection and annotation, especially in the cases of cross-domain learning. One way to circumvent this cost is constructing models to synthesize data and provide automatic annotation. Although these models are attractive, they often can not be generalized from synthetic images to real-world images. Therefore, domain adaptive algorithm is needed to improve these models, so that they can be applied successfully. In this paper, we propose a novel unsupervised domain adaptive framework codenamed D-DANN inspired by the theory of adversarial learning. We apply the discriminator to diverse the features extracted from dual branch CNN. We can obtain more sufficient shared representation across domains by the proposed dual feature extractors. The framework can be easily adapt to most popular CNN models to improve the representation power. We implement the D-DANN with several popular CNN models including LeNet, AlexNet and so on. Using these D-DANN enhanced neural networks, we conduct extensive experiments on several pairs of domain adaptive validation datasets. The results show that our approach can efficiently enhance domain adaptive capability of general CNN models for unlabeled data.
Yuchun Fang, Qiulong Yuan, Wei Zhang 0021, Zhaoxiang Zhang 0001
ICPR3
2018 UAV Target Tracking with A Boundary-Decision Network
abstract
The aspect ratio of a target changes frequently during UAV tracking task, which makes the aerial tracking very challenging. Traditional trackers struggle from such problem as they mainly focus on the scale variation issue by maintaining a certain aspect ratio. In this paper, we propose a novel tracker, named boundary-decision network (BDNet), to address the aspect ratio variation in UAV tracking. Unlike previous work, the proposed method aims at operating each boundary separately with a policy network. Given an initial estimate of the bounding box, a sequential actions are generated to tune the four boundaries with an optimization strategy including boundary proposal rejection, offline and online learning. Experimental results on the benchmark aerial dataset prove that the proposed approach outperforms existing trackers and produces significant accuracy gains in dealing with the aspect ratio variation in UAV tracking.
Ke Song 0003, Wei Zhang 0021, Xuewen Rong
ICPR2
2018 Image-level to Pixel-wise Labeling: From Theory to Practice
abstract
Conventional convolutional neural networks (CNNs) have achieved great success in image semantic segmentation. Existing methods mainly focus on learning pixel-wise labels from an image directly. In this paper, we advocate tackling the pixel-wise segmentation problem by considering the image-level classification labels. Theoretically, we analyze and discuss the effects of image-level labels on pixel-wise segmentation from the perspective of information theory. In practice, an end-to-end segmentation model is built by fusing the image-level and pixel-wise labeling networks. A generative network is included to reconstruct the input image and further boost the segmentation model training with an auxiliary loss. Extensive experimental results on benchmark dataset demonstrate the effectiveness of the proposed method, where good image-level labels can significantly improve the pixel-wise segmentation accuracy.
Tiezhu Sun, Wei Zhang 0021, Zhijie Wang 0010, Lin Ma 0002, Zequn Jie
IJCAI2
2018 Master-Slave Curriculum Design for Reinforcement Learning
abstract
Curriculum learning is often introduced as a leverage to improve the agent training for complex tasks, where the goal is to generate a sequence of easier subasks for an agent to train on, such that final performance or learning speed is improved. However, conventional curriculum is mainly designed for one agent with fixed action space and sequential simple-to-hard training manner. Instead, we present a novel curriculum learning strategy by introducing the concept of master-slave agents and enabling flexible action setting for agent training. Multiple agents, referred as master agent for the target task and slave agents for the subtasks, are trained concurrently within different action spaces by sharing a perception network with an asynchronous strategy. Extensive evaluation on the VizDoom platform demonstrates the joint learning of master agent and slave agents mutually benefit each other. Significant improvement is obtained over A3C in terms of learning speed and performance.
Yuechen Wu, Wei Zhang 0021, Ke Song 0003
IJCAI2
2018 SCRATCH: A Scalable Discrete Matrix Factorization Hashing for Cross-Modal Retrieval
abstract
In recent years, many hashing methods have been proposed for the cross-modal retrieval task. However, there are still some issues that need to be further explored. For example, some of them relax the binary constraints to generate the hash codes, which may generate large quantization error. Although some discrete schemes have been proposed, most of them are time-consuming. In addition, most of the existing supervised hashing methods use an n x n similarity matrix during the optimization, making them unscalable. To address these issues, in this paper, we present a novel supervised cross-modal hashing method---Scalable disCRete mATrix faCtorization Hashing, SCRATCH for short. It leverages the collective matrix factorization on the kernelized features and the semantic embedding with labels to find a latent semantic space to preserve the intra- and inter-modality similarities. In addition, it incorporates the label matrix instead of the similarity matrix into the loss function. Based on the proposed loss function and the iterative optimization algorithm, it can learn the hash functions and binary codes simultaneously. Moreover, the binary codes can be generated discretely, reducing the quantization error generated by the relaxation scheme. Its time complexity is linear to the size of the dataset, making it scalable to large-scale datasets. Extensive experiments on three benchmark datasets, namely, Wiki, MIRFlickr-25K, and NUS-WIDE, have verified that our proposed SCRATCH model outperforms several state-of-the-art unsupervised and supervised hashing methods for cross-modal retrieval.
Chuan-Xiang Li, Zhen-Duo Chen 0001, Peng-Fei Zhang 0001, Xin Luo 0006, Liqiang Nie, Wei Zhang 0021, Xin-Shun Xu
ACM Multimedia6
2018 Emotion recognition by assisted learning with convolutional neural networks
Xuanyu He, Wei Zhang 0021
Neurocomputing2
2018 Long-range terrain perception using convolutional neural networks
Wei Zhang 0021, Qi Chen 0005, Weidong Zhang 0005, Xuanyu He
Neurocomputing1
2018 Learning Bidirectional Temporal Cues for Video-Based Person Re-Identification
abstract
This paper presents an end-to-end learning architecture for video-based person re-identification by integrating convolutional neural networks (CNNs) and bidirectional recurrent neural networks (BRNNs). Given a video with consecutive frames, features of each frame are extracted with CNN and then are fed into the BRNN to get a final spatio-temporal representation about the video. Specifically, CNN acts as a Spatial Feature Extractor, while BRNN is expected to capture the temporal cues of sequential frames in both forward and backward directions, simultaneously. The whole network is trained end-to-end with a joint identification and verification manner. Experimental results on benchmark data sets show that the proposed model can effectively learn spatio-temporal features relevant for re-identification and outperforms existing video-based person re-identification methods.
Wei Zhang 0021, Xuanyu He
IEEE Trans. Circuits Syst. Video Technol.1
2018 A Feature Descriptor Based on Local Normalized Difference for Real-World Texture Classification
abstract
In this paper, we propose a normalized difference vector (NDV) for texture representation. Compared to local-binary-pattern-based descriptors, the proposed NDV takes full advantage of the local difference, and the size can be extended flexibly to cover a large local region. We further employ the bag-of-words model to integrate the local descriptors into a global feature representation of an image. In addition, two strategies are introduced for the proposed NDV to achieve rotation invariance. We test the proposed texture descriptor on benchmark datasets, such as AniTex, VehApp, KTH-TIPS2a, OpenSurface, and Kylberg. Classification results demonstrate the superiority of the proposed descriptor over state-of-the-art methods.
Wei Zhang 0021, Weidong Zhang 0005, Kan Liu 0001, Jason Gu
IEEE Trans. Multim.1
2017 Exploiting patch-based correlation for ghost removal in exposure fusion
abstract
In this paper, we present a robust exposure fusion algorithm to tackle the problems of motion removal and detail preserving in dynamic scenes. With one exposure as reference, the motion appeared in the exposure stack can be detected by comparing the structural consistency, which is extracted by measuring the degree of linear correlation between the patches of the reference image and the other source images. Then, a stack of latent images with consistent contents can be synthesized after motion removal. For detail preserving, a contrast criterion is introduced to measure the exposedness and generate visibility maps of each latent image. Guided by the visibility maps, a tonemapped-like HDR image which is ghost-free and with all details preserved could be produced by seamlessly merging the latent images. Exposure fusion tests on various dynamic scenes demonstrate the superiority of the proposed method over existing state-of-the-art approaches.
Shengnan Hu, Wei Zhang 0021
ICME2
2017 SIFT-based adaptive prediction structure for light field compression
abstract
A light field consists of multiple views of a scene, which can be arranged and encoded like a pseudo sequence. Since the correlations between views are not equal and indeed content dependent, a well-constructed coding order and adaptive prediction structure will improve performance. In this paper, we propose an adaptive prediction structure for light field compression. While the coding order is inherited from the 2-D hierarchical coding order, the prediction structure is determined by the differences between scale-invariant feature transform (SIFT) descriptors of the views. Experimental results show that the proposed method leads to on average 5.71% BD-rate reduction compared with fixed prediction structure.
Wei Zhang 0021, Dong Liu 0002, Zhiwei Xiong, Jizheng Xu
VCIP1
2017 Patch-Based correlation for deghosting in exposure fusion
Wei Zhang 0021, Shengnan Hu, Kan Liu 0001
Inf. Sci.1
2017 Motion-free exposure fusion based on inter-consistency and intra-consistency
Wei Zhang 0021, Shengnan Hu, Kan Liu 0001
Inf. Sci.1
2017 Optimal seamline detection in dynamic scenes via graph cuts for image mosaicking
Li Li 0047, Jian Yao 0002, Haoang Li, Menghan Xia, Wei Zhang 0021
Mach. Vis. Appl.5
2017 Globally consistent alignment for planar mosaicking via topology analysis
Menghan Xia, Jian Yao 0002, Renping Xie, Li Li 0047, Wei Zhang 0021
Pattern Recognit.5
2017 Video-Based Pedestrian Re-Identification by Adaptive Spatio-Temporal Appearance Model
abstract
Pedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts independently. However, a body part can also appear differently during different phases of an action. In this paper, we consider the temporal alignment problem, in addition to the spatial one, and propose a new approach that takes the video of a walking person as input and builds a spatiotemporal appearance representation for pedestrian re-identification. Particularly, given a video sequence, we exploit the periodicity exhibited by a walking person to generate a spatiotemporal body-action model, which consists of a series of body-action units corresponding to certain action primitives of certain body parts. Fisher vectors are learned and extracted from individual body-action units and concatenated into the final representation of the walking person. Unlike previous spatiotemporal features that only take into account local dynamic appearance information, our representation aligns the spatiotemporal appearance of a pedestrian globally. Extensive experiments on public data sets show the effectiveness of our approach compared with the state of the art.
Wei Zhang 0021, Bingpeng Ma, Kan Liu 0001, Rui Huang 0001
IEEE Trans. Image Process.1
2017 Learning to Predict High-Quality Edge Maps for Room Layout Estimation
abstract
The goal of room layout estimation is to predict the three-dimensional box that represents the room spatial structure from a monocular image. In this paper, a deconvolution network is trained first to predict the edge map of a room image. Compared to the previous fully convolutional networks, the proposed deconvolution network has a multilayer deconvolution process that can refine the edge map estimate layer by layer. The deconvolution network also has fully connected layers to aggregate the information of every region throughout the entire image. During the layout generation process, an adaptive sampling strategy is introduced based on the obtained high-quality edge maps. Experimental results prove that the learned edge maps are highly reliable and can produce accurate layouts of room images.
Weidong Zhang 0005, Wei Zhang 0021, Kan Liu 0001, Jason Gu
IEEE Trans. Multim.2
2016 Edge chain detection by applying Helmholtz principle on gradient magnitude map
abstract
In this paper, we present an efficient edge chain detection algorithm by applying the Helmholtz principle on the gradient magnitude map of an image. An edge chain validation method is proposed which uses the “relative number of false alarms” (RNFA) instead of the traditional “number of false alarms” (NFA). The edge chains are detected first and then validated according to their RNFA values. In this way, edge chains that are weak in gradients but meaningful in vision can be detected. To evaluate the proposed edge chain detector in quantity, an edge chain detection benchmark which consists of 25 labeled images in different scenes was built. The proposed edge chain detector was tested in this benchmark, and the experimental results sufficiently demonstrate that the proposed edge chain detector outperforms the state-of-the-art methods.
Xiaohu Lu, Jian Yao 0002, Li Li 0047, Wei Zhang 0021
ICPR5
2016 Pyramid stereo matching for spherical panoramas
abstract
This paper presents a novel pyramid stereo matching method to improve the matching accuracy of panoramas. Initial camera parameters and feature correspondences are obtained from Structure From Motion (SFM) with normal images extracted from two panoramas. Then a stereo matching pyramid is constructed to refine the feature correspondences layer by layer, and the correspondence is corrected in the original panoramas. Experimental results show the matching accuracy gains provided by the proposed approach.
Jian Weng 0007, Wei Zhang 0021, Weidong Zhang 0005, Jianjie Gao
VCIP2
2016 Hybrid human detection and recognition in surveillance
Qiang Liu 0015, Wei Zhang 0021, Hongliang Li 0001, King Ngi Ngan
Neurocomputing2
2016 Deep Neural Networks for wireless localization in indoor and outdoor environments
Wei Zhang 0021, Kan Liu 0001, Weidong Zhang 0005, Youmei Zhang, Jason Gu
Neurocomputing1
2016 Learning structure of stereoscopic image for no-reference quality assessment with convolutional neural network
Wei Zhang 0021, Chenfei Qu, Lin Ma 0002, Jingwei Guan, Rui Huang 0001
Pattern Recognit.1
2015 A Spatio-Temporal Appearance Representation for Viceo-Based Pedestrian Re-Identification
abstract
Pedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts independently. However, a body part can also appear differently during different phases of an action. In this paper we consider the temporal alignment problem, in addition to the spatial one, and propose a new approach that takes the video of a walking person as input and builds a spatio-temporal appearance representation for pedestrian re-identification. Particularly, given a video sequence we exploit the periodicity exhibited by a walking person to generate a spatio-temporal body-action model, which consists of a series of body-action units corresponding to certain action primitives of certain body parts. Fisher vectors are learned and extracted from individual body-action units and concatenated into the final representation of the walking person. Unlike previous spatio-temporal features that only take into account local dynamic appearance information, our representation aligns the spatio-temporal appearance of a pedestrian globally. Extensive experiments on public datasets show the effectiveness of our approach compared with the state of the art.
Kan Liu 0001, Bingpeng Ma, Wei Zhang 0021, Rui Huang 0001
ICCV3
2015 No-reference blur assessment based on edge modeling
Jingwei Guan, Wei Zhang 0021, Jason Gu, Hongliang Ren 0001
J. Vis. Commun. Image Represent.2
2015 Multimodal learning for facial expression recognition
Wei Zhang 0021, Youmei Zhang, Lin Ma 0002, Jingwei Guan, Shijie Gong
Pattern Recognit.1
2014 Computer-Aided Bleeding Detection in WCE Video
abstract
Wireless capsule endoscopy (WCE) can directly take digital images in the gastrointestinal tract of a patient. It has opened a new chapter in small intestine examination. However, a major problem associated with this technology is that too many images need to be manually examined by clinicians. Currently, there is no standard for capsule endoscopy image interpretation and classification. Most state-of-the-art CAD methods often suffer from poor performance, high computational cost, or multiple empirical thresholds. In this paper, a new method for rapid bleeding detection in the WCE video is proposed. We group pixels through superpixel segmentation to reduce the computational complexity while maintaining high diagnostic accuracy. Feature of each superpixel is extracted using the red ratio in RGB space and fed into support vector machine for classification. Also, the influence of edge pixels has been removed in this paper. Comparative experiments show that our algorithm is superior to the existing methods in terms of sensitivity, specificity, and accuracy.
Yanan Fu, Wei Zhang 0021, Mrinal Mandal 0001, Max Q.-H. Meng
IEEE J. Biomed. Health Informatics2
2012 Reference-guided exposure fusion in dynamic scenes
Wei Zhang 0021, Wai-kuen Cham
J. Vis. Commun. Image Represent.1
2012 Single-Image Refocusing and Defocusing
abstract
In this paper, we present a postprocessing method to tackle the single-image refocusing-and-defocusing problem. The proposed method can accomplish the tasks of focus-map estimation and image refocusing and defocusing. Given an image with a mixture of focused and defocused objects, we first detect the edges and then estimate the focus map based on the edge blurriness, which is depicted explicitly by a parametric model. The image refocusing problem is addressed in a blind deconvolution framework, where the image prior is modeled by using both global and local constraints. In particular, we correct the defocused blurry edges to sharp ones with the aid of the parametric edge model and then render this cue as a local prior to ensure the sharpness of the refocused image. Experimental results demonstrate that the proposed method performs well in producing visually plausible images with different focus effects from a single input.
Wei Zhang 0021, Wai-kuen Cham
IEEE Trans. Image Process.1
2012 Gradient-Directed Multiexposure Composition
abstract
In this paper, we present a simple yet effective method that takes advantage of the gradient information to accomplish the multiexposure image composition in both static and dynamic scenes. Given multiple images with different exposures, the proposed approach is capable of producing a pleasant tone-mapped-like high-dynamic-range image by compositing them seamlessly with the guidance of gradient-based quality assessment. In particular, two novel quality measures, namely, visibility and consistency, are developed based on the observations of gradient changes among different exposures. Experiments in various static and dynamic scenes are conducted to demonstrate the effectiveness of the proposed method.
Wei Zhang 0021, Wai-kuen Cham
IEEE Trans. Image Process.1
2011 Hallucinating Face in the DCT Domain
abstract
In this paper, we propose a novel learning-based face hallucination framework built in the DCT domain, which can produce a high-resolution face image from a single low-resolution one. The problem is formulated as inferring the DCT coefficients in frequency domain instead of estimating pixel intensities in spatial domain. Our study shows that DC coefficients can be estimated fairly accurately by simple interpolation-based methods. AC coefficients, which contain the information of local features of face image, cannot be estimated well using interpolation. A simple but effective learning-based inference model is proposed to infer the ac coefficients. Experiments have been conducted to demonstrate the effectiveness of the proposed method in producing high quality hallucinated face images.
Wei Zhang 0021, Wai-kuen Cham
IEEE Trans. Image Process.1
2010 Gradient-directed composition of multi-exposure images
abstract
In this paper, we present a simple yet effective method that takes advantage of the gradient information to accomplish the multi-exposure image composition in both static and dynamic scenes. Given multiple images with different exposures, the proposed approach is capable of producing a pleasant tonemapped-like high dynamic range (HDR) image by compositing them seamlessly with the guidance of gradient-based quality assessment. Especially, two novel quality measures: visibility and consistency, are developed based on the observations of gradient changes among different exposures. Experiments in various static and dynamic scenes are conducted to demonstrate the effectiveness of the proposed method.
Wei Zhang 0021, Wai-kuen Cham
CVPR1
2010 High quality artifact-free super-resolution
abstract
Blurring and jaggy artifacts are the primal culprits that plague the current super-resolution techniques. In this paper, we propose a simple but effective approach which is capable of producing a pleasant artifact-free high-resolution image from a single low-resolution input. Specifically, we first magnify the low-resolution image to the desired resolution through structure adaptive interpolation to avoid jaggies. Then a salient edge directed deblurring scheme is introduced to remove the blurriness of the magnified image. Unlike previous work, we advocate solving the deblurring problem in an efficient manner with the aid of little user intervention. Our study shows that the blurring kernel can be estimated fairly well from the salient edges selected with user-drawn stroke based on a parametric edge model. Experiments are conducted to validate the effectiveness of the proposed method.
Wei Zhang 0021, Wai-kuen Cham
ICIP1
2010 3D Modeling from Multiple Images
Wei Zhang 0021, Jian Yao 0001, Wai-kuen Cham
ISNN (2)1
2008 Learning-based face hallucination in DCT domain
abstract
In this paper, we propose a novel learning-based face hallucination framework built in DCT domain, which can recover the high-resolution face image from a single low-resolution one. Unlike most previous learning-based work, our approach addresses the face hallucination problem from a different angle. In details, the problem is formulated as inferring DCT coefficients in frequency domain instead of estimating pixel intensities in spatial domain. Experimental results show that DC coefficients can be estimated fairly accurately by simple interpolation-based methods. AC coefficients, which contain the information of local features of face image, cannot be estimated well using interpolation. We propose a method to infer AC coefficients by introducing an efficient learning-based inference model. Moreover, the proposed framework can lead to significant savings in memory and computation cost since the redundancy of the training set is reduced a lot by clustering. Experimental results demonstrate that our approach is very effective to produce hallucinated face images with high quality.
Wei Zhang 0021, Wai-kuen Cham
CVPR1
2008 A single image based blind super-resolution approach
abstract
In this paper, we address the problem of producing super-resolved image from a single low-resolution input. Unlike most previous work, the camera’s Point Spread Function (PSF) is not assumed to be known in advance and the single image super-resolution problem is formulated as a blind deconvolution problem under a MAP framework which can be optimized effectively in an iterative manner. Experimental results demonstrate that our method can successfully generate super-resolved image with high quality, both subjectively and objectively.
Wei Zhang 0021, Wai-kuen Cham
ICIP1
2004 A stereo matching algorithm based on multiresolution and epipolar constraint
abstract
Stereo matching is one of the most active research areas in computer vision. In this paper, a fast stereo matching algorithm by means of epipolar constraint and multiresolution approach was presented. The searching scope of corresponding pixels in the original image is obtained and has diminished a lot based on multiresolution approach. Then intensity correlation principle and epipolar constraint can be applied to get the stereo matching results in this scope. In this way, we reduce the search time for correspondence and ensure the validity of matching. The experimental results show this algorithm is effective and efficient.
Wei Zhang 0021, Quanbing Zhang, Sui Wei
ICIG1