Zhiyong Liu 0001

dblp:16/5205-1 · also Zhi-Yong Liu 0001 · DBLP profile ↗
← Back
78ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0003-2148-1846ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 66 · 10 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Fine-Grained Multimodal Alignment for Image-Text Retrieval via Graph Learning
Mao Chen 0007, Xiangkai Zhang, Lu Qi 0001, Xiangtai Li, Xu Yang 0004, Steven C. H. Hoi, Zhiyong Liu 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.7
2026 AutoTraj: Autoregressive trajectory synthesis for OOD task adaptation in offline meta-RL
Zhenyang Lin, Yurou Chen, Zhiyong Liu 0001
Neurocomputing4
2025 CARD: Control-Driven Autoregressive Reconstruction with Decoupled Learning for Multi-Class Anomaly Detection
abstract
Multi-class unsupervised anomaly detection (UAD) is challenging due to the difficulty of harmonizing distributional differences across categories within a unified framework. While recent diffusion-based methods have demonstrated promising performance by leveraging denoising processes for anomaly reconstruction, the lack of explicit causal constraints limits their ability to handle complex logical inconsistencies. Inspired by the success of autoregressive models in enforcing local-to-global logical consistency, we propose a Control-driven Autoregressive Reconstruction with Decoupled learning (CARD) framework for multi-class UAD. It first tokenizes images using vector quantization and employs a vision autoregressive model to capture causal dependencies within normal patterns as explicit prior knowledge. Then, we introduce a Control-Driven Reconstruction (CDR) network, which aligns input features as control signals into the frozen autoregressive model to adjust the predicted distribution, enabling a generative reconstruction process. Additionally, we apply perturbations to the CDR input to simulate anomalous conditions, facilitating the model to correct out-of-distribution anomaly features under prior knowledge constraints. By decoupling normal pattern learning from reconstruction, CARD prevents identity mapping caused by forgetting implicit priors in the conventional reconstruction-based method. Comprehensive experiments on several benchmark datasets validate the effectiveness of our approach. On the MVTecAD dataset, CARD achieved AUROC scores of 98.7% at the image level and 98.5% at the pixel level.
Mingqing Wang, Boyi Sun, Qianfan Zhao, Lu Zhang 0054, Zhiyong Liu 0001, Xu Yang 0004, Suiwu Zheng
MMAsia6
2025 Adaptive Video-Conditioned Imitation Learning via Bidirectional Cross-Domain Skill Transfer
abstract
Imitation learning by watching humans offers a promising path to learning general-purpose robot skills with intuitive task specifications. While prior approaches for video-conditioned imitation learning directly extract skill embeddings from unstructured human videos and follow demonstrations step-by-step, such paradigm usually falls short in generalizing to unseen long-horizon tasks with a single human prompt video due to the significant embodiment and environment gap. To this end, our key insight is to infer local intentions from videos in order to retrieve robot skill memories from prior experience, and conversely select the feasible video clip to follow based on robot observations. Motivated by this, we introduce AdaMimic, a hierarchical imitation learning method that learns the bidirectional mapping of cross-domain sensorimotor skills and derives skill-based policy conditioned on adaptable latent plans. To enable generalization to unseen tasks given cross-domain human videos, AdaMimic leverages task-agnostic play data for interaction-aware skill embedding extraction and video-robot trajectory pairs for semantic and temporal human-to-robot skill alignment. In addition, our method exploits a skill adapter for robot-to-human alignment to adaptively align the robot with the skill intentions. We systematically evaluate AdaMimic on both simulated and real-world kitchen domains, demonstrating AdaMimic’s superiority over prior imitation learning methods in generalizing to novel long-horizon tasks with a single human prompt video.
Zhenyang Lin, Yurou Chen, Xianxiang Zhang, Bin Liang 0001, Zhiyong Liu 0001
IEEE Trans Autom. Sci. Eng.6
2025 DeepPartitioning: Deep Learning of Graph Partitioning for Neuron Segmentation From Electron Microscopy Volume via Graph Neural Network
abstract
Superpixel aggregation represents a highly effective approach for automated neuron segmentation from electron microscopy (EM) volumes, which can be considered as a graph partitioning task on the region adjacency graph (RAG) of extracted superpixels. However, existing graph partitioning models for superpixel aggregation suffer from the modeling error due to insufficient model capacity. More specifically, the modeling error is caused by the simplification in formulating the real-world graph partitioning task (i.e., superpixel aggregation) into a mathematically well-defined optimization problem. To address this issue, we sidestep the explicit formulation and propose a fully end-to-end superpixel aggregation method based on deep learning of the graph partitioning task, called DeepPartitioning. The central challenge lies in characterizing the partitioning task involving combinatorial complexity. Hence, our method incorporates a line graph neural network (LGNN) to capture higher-order relational structures in RAGs. Specifically, the LGNN enables the propagation of second-order superpixel-pair features among adjacent edges in RAGs. In this way, the partitioning task can be implicitly transformed into the vanilla second-order multicut problem while maintaining higher-order structural information. Overall, our method integrates a second-order feature extractor, a higher-order feature integrator (i.e., the LGNN), and a differentiable approximation to a multicut solver into a unified, learnable framework. Extensive experiments on three public EM datasets demonstrate the effectiveness of the proposed DeepPartitioning within the neuron segmentation pipeline.
Zhenchen Li, Xu Yang 0004, Jing Liu 0054, Zhiyong Liu 0001, Hua Han 0001
IEEE Trans. Medical Imaging7
2025 Deep Graph Reinforcement Learning for Solving Multicut Problem
abstract
The multicut problem, also known as correlation clustering, is a classic combinatorial optimization problem that aims to optimize graph partitioning given only node (dis)similarities on edges. It serves as an elegant generalization for several graph partitioning problems and has found successful applications in various areas such as data mining and computer vision. However, the multicut problem with an exponentially large number of cycle constraints proves to be NP-hard, and existing solvers either suffer from exponential complexity or often give unsatisfactory solutions due to inflexible heuristics driven by hand-designed mechanisms. In this article, we propose a deep graph reinforcement learning method to solve the multicut problem within a combinatorial decision framework involving sequential edge contractions. The customized subgraph neural network adapts to the dynamically edge-contracted graph environment by extracting bilevel connected features from both contracted and original graphs. Our method can learn to infer feasible multicut solutions end-to-end toward optimization of the multicut objective in a data-driven manner. More specifically, by exploring the decision space adaptively, it implicitly gains heuristic knowledge from topological patterns of instances and thereby generates more targeted heuristics overcoming the short-sightedness inherent in the hand-designed ones. During testing, the learned heuristics iteratively contract graphs to construct high-quality solutions within polynomial time. Extensive experiments on synthetic and real-world multicut instances show the superiority of our method over existing combinatorial solvers, while also maintaining a certain level of out-of-distribution generalization ability.
Zhenchen Li, Xu Yang 0004, Shaofeng Zeng, Jingbin Yuan, Zhiyong Liu 0001, Hua Han 0001
IEEE Trans. Neural Networks Learn. Syst.7
2025 Weakly Aligned Feature Fusion for Multimodal Object Detection
abstract
To achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image pair is not strictly aligned, making one object has different positions in different modalities. For the deep learning method, this problem makes it difficult to fuse multimodal features and puzzles the convolutional neural network (CNN) training. In this article, we propose a general multimodal detector named aligned region CNN (AR-CNN) to tackle the position shift problem. First, a region feature (RF) alignment module with adjacent similarity constraint is designed to consistently predict the position shift between two modalities and adaptively align the cross-modal RFs. Second, we propose a novel region of interest (RoI) jitter strategy to improve the robustness to unexpected shift patterns. Third, we present a new multimodal feature fusion method that selects the more reliable feature and suppresses the less useful one via feature reweighting. In addition, by locating bounding boxes in both modalities and building their relationships, we provide novel multimodal labeling named KAIST-Paired. Extensive experiments on 2-D and 3-D object detection, RGB-T, and RGB-D datasets demonstrate the effectiveness and robustness of our method.
Lu Zhang 0054, Zhiyong Liu 0001, Xiangyu Zhu 0001, Zhan Song, Xu Yang 0004, Zhen Lei 0001, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Hierarchical Human-to-Robot Imitation Learning for Long-Horizon Tasks via Cross-Domain Skill Alignment
abstract
For a general-purpose robot, it is desirable to imitate human demonstration videos that can effectively solve long-horizon tasks and perform novel ones. Recent advances in skill-based imitation learning have shown that extracting skill embedding from raw human videos is a promising paradigm to enable robots to cope with long-horizon tasks. However, generalization to unseen tasks in a different domain with a human prompt video poses a significant challenge due to the big embodiment and environment difference. To this end, we present Hierarchical Human-to-Robot Imitation Learning (H2RIL) that learns the mapping of cross-domain sensorimotor skills and utilizes it to generalize to unseen tasks given a human video in a different environment. To allow for generalizing zero-shot across environments and embodiments, H2RIL leverages task-agnostic play data for low-level policy training and paired human-robot data for both semantic and temporal skill embedding alignment. Extensive experiments in a simulated kitchen environment demonstrate that H2RIL significantly outperforms other prior baselines and is capable of generalizing to composable new tasks and adapting to Out-of-Distribution (OOD) tasks.
Zhenyang Lin, Yurou Chen, Zhiyong Liu 0001
ICRA3
2024 Efficient Offline Meta-Reinforcement Learning via Robust Task Representations and Adaptive Policy Generation
Zhenyang Lin, Yurou Chen, Zhiyong Liu 0001
IJCAI4
2024 Enhancing class-incremental object detection in remote sensing through instance-aware distillation
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001
Neurocomputing4
2024 Active domain adaptation for semantic segmentation via dynamically balancing domainness and uncertainty
Lu Zhang 0054, Zhiyong Liu 0001
Image Vis. Comput.3
2024 Adaptive bias-variance trade-off in advantage estimator for actor-critic algorithms
Yurou Chen, Zhiyong Liu 0001
Neural Networks3
2024 DeepMulticut: Deep Learning of Multicut Problem for Neuron Segmentation From Electron Microscopy Volume
abstract
Superpixel aggregation is a powerful tool for automated neuron segmentation from electron microscopy (EM) volume. However, existing graph partitioning methods for superpixel aggregation still involve two separate stages-model estimation and model solving, and therefore model error is inherent. To address this issue, we integrate the two stages and propose an end-to-end aggregation framework based on deep learning of the minimum cost multicut problem called DeepMulticut. The core challenge lies in differentiating the NP-hard multicut problem, whose constraint number is exponential in the problem size. With this in mind, we resort to relaxing the combinatorial solver-the greedy additive edge contraction (GAEC)-to a continuous Soft-GAEC algorithm, whose limit is shown to be the vanilla GAEC. Such relaxation thus allows the DeepMulticut to integrate edge cost estimators, Edge-CNNs, into a differentiable multicut optimization system and allows a decision-oriented loss to feed decision quality back to the Edge-CNNs for adaptive discriminative feature learning. Hence, the model estimators, Edge-CNNs, can be trained to improve partitioning decisions directly while beyond the NP-hardness. Also, we explain the rationale behind the DeepMulticut framework from the perspective of bi-level optimization. Extensive experiments on three public EM datasets demonstrate the effectiveness of the proposed DeepMulticut.
Zhenchen Li, Xu Yang 0004, Bei Hong, Hao Zhai 0003, Lijun Shen, Xi Chen 0031, Zhiyong Liu 0001, Hua Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 Automatically Discovering Novel Visual Categories With Adaptive Prototype Learning
abstract
This article targets the task of novel category discovery (NCD), which aims to discover unknown categories when a certain number of classes are already known. The NCD task is challenging due to its closeness to real-world scenarios, where we have only encountered some partial classes and corresponding images. Unlike previous approaches to NCD, we propose a novel adaptive prototype learning method that leverages prototypes to emphasize category discrimination and alleviate the issue of missing annotations for novel classes. Concretely, the proposed method consists of two main stages: prototypical representation learning and prototypical self-training. In the first stage, we develop a robust feature extractor that could effectively handle images from both base and novel categories. This ability of instance and category discrimination of the feature extractor is boosted by self-supervised learning and adaptive prototypes. In the second stage, we utilize the prototypes again to rectify offline pseudo labels and train a final parametric classifier for category clustering. We conduct extensive experiments on four benchmark datasets, demonstrating our method's effectiveness and robustness with state-of-the-art performance.
Lu Zhang 0054, Lu Qi 0001, Xu Yang 0004, Hong Qiao, Ming-Hsuan Yang 0001, Zhiyong Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Refined Pseudo Labeling for Source-Free Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) assumes that both labeled source data and unlabeled target data are available for training, but this assumption does not always hold in real-world scenarios. Thus, source-free DAOD is proposed to adapt the source-trained detectors to target domains with only unlabeled target data. Existing source-free DAOD methods typically utilize pseudo labeling, where the performance heavily relies on the selection of confidence threshold. However, most prior works adopt a single fixed threshold for all classes to generate pseudo labels, which ignore the imbalanced class distribution, resulting in biased pseudo labels. In this work, we propose a refined pseudo labeling framework for source-free DAOD. First, to generate unbiased pseudo labels, we present a category-aware adaptive threshold estimation module, which adaptively provides the appropriate threshold for each category. Second, to alleviate incorrect box regression, a localization-aware pseudo label assignment strategy is introduced to divide labels into certain and uncertain ones and optimize them separately. Finally, extensive experiments on four adaptation tasks demonstrate the effectiveness of our method.
Lu Zhang 0054, Zhiyong Liu 0001
ICASSP3
2023 Unseen Object Instance Segmentation with Fully Test-time RGB-D Embeddings Adaptation
abstract
Segmenting unseen objects is a crucial ability for the robot since it may encounter new environments during the operation. Recently, a popular solution is leveraging RGB-D features of large-scale synthetic data and directly applying the model to unseen real-world scenarios. However, the domain shift caused by the sim2real gap is inevitable, posing a crucial challenge to the segmentation model. In this paper, we em-phasize the adaptation process across sim2real domains and model it as a learning problem on the BatchNorm param-eters of a simulation-trained model. Specifically, we propose a novel non-parametric entropy objective, which formulates the learning objective for the test-time adaptation in an open-world manner. Then, a cross-modality knowledge distillation objective is further designed to encourage the test-time knowledge transfer for feature enhancement. Our approach can be efficiently implemented with only test images, without requiring annotations or revisiting the large-scale synthetic training data. Besides significant time savings, the proposed method consistently improves segmentation results on the overlap and boundary metrics, achieving state-of-the-art performance on unseen object instance segmentation.
Lu Zhang 0054, Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001
ICRA5
2023 Zero-Shot Object Goal Visual Navigation
abstract
Object goal visual navigation is a challenging task that aims to guide a robot to find the target object based on its visual observation, and the target is limited to the classes pre-defined in the training stage. However, in real households, there may exist numerous target classes that the robot needs to deal with, and it is hard for all of these classes to be contained in the training stage. To address this challenge, we study the zero-shot object goal visual navigation task, which aims at guiding robots to find targets belonging to novel classes without any training samples. To this end, we also propose a novel zero-shot object navigation framework called semantic similarity network (SSNet). Our framework use the detection results and the cosine similarity between semantic word embeddings as input. Such type of input data has a weak correlation with classes and thus our framework has the ability to generalize the policy to novel classes. Extensive experiments on the AI2-THOR platform show that our model outperforms the baseline models in the zero-shot object navigation task, which proves the generalization ability of our model. Our code is available at: https://github.com/pioneer-innovation/Zero-Shot-Object-Navigation.
Qianfan Zhao, Lu Zhang 0054, Bin He 0003, Hong Qiao, Zhiyong Liu 0001
ICRA5
2023 Incremental Few-Shot Object Detection with scale- and centerness-aware weight generation
Lu Zhang 0054, Xu Yang 0004, Lu Qi 0001, Shaofeng Zeng, Zhiyong Liu 0001
Comput. Vis. Image Underst.5
2023 Frequency-based pseudo-domain generation for domain generalizable object detection
Lu Zhang 0054, Zhiyong Liu 0001
Neurocomputing3
2023 RTDOD: A large-scale RGB-thermal domain-incremental object detection dataset for UAVs
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001
Image Vis. Comput.6
2022 FIT: Frequency-Based Image Translation for Domain Adaptive Object Detection
Lu Zhang 0054, Zhiyong Liu 0001, Hangtao Feng
ICONIP (3)3
2022 A Learning Method for Feature Correspondence with Outliers
abstract
Feature correspondence is an important topic in many computer vision or robot vision tasks. Different from traditional optimization based matching method, in the last two years, researchers are finally able to solve the matching process in a learning manner. As a representative method, SuperGlue achieves superior performance in many real-world tasks, but it still has problems in dealing with outlier features. Targeting at the outlier problem, this paper improves SuperGlue by introducing a deep learning based feature correspondence method, which consists of the pruned attentional graph neural network and the improved matching layer for the outlier problem. Experiments on real world images validate the effectiveness of the proposed method.
Xu Yang 0004, Shaofeng Zeng, Yu-Chen Lu, Zhiyong Liu 0001
ICPR5
2022 Safety-based Reinforcement Learning Longitudinal Decision for Autonomous Driving in Crosswalk Scenarios
abstract
Autonomous vehicles (AVs) need to make driving decisions to interact with other traffic participants. By adapting to different scenarios with specific parameters, traditional strategies attempt to leverage rule-based methods to solve the decision problems. In this paper, we present a novel reinforcement learning method for resolving interaction uncertainty in the decision-making problem. We construct prior knowledge by introducing traffic regulations and constraints and then converting them into rules that govern the learning of driving policies. To promote safe driving, a safety-aware module equipped with a mathematical collision correlation analysis is developed to anticipate and handle dangerous traffic scenarios. A realistic scenario involving an AV approaching a crosswalk is used to validate the proposed method. The experimental results indicate that the proposed method improves driving safety and efficiency significantly when compared to alternative approaches and can be generalized to more difficult scenarios.
Fangzhou Xiong, Dongchun Ren, Mingyu Fan, Shuguang Ding, Zhiyong Liu 0001
IJCNN5
2022 Markovian policy network for efficient robot learning
Yurou Chen, Zhiyong Liu 0001
Neurocomputing3
2022 Graph matching based on fast normalized cut and multiplicative update mapping
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001
Pattern Recognit.4
2022 Incremental few-shot object detection via knowledge transfer
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001
Pattern Recognit. Lett.4
2021 Adaptive Advantage Estimation for Actor-Critic Algorithms
abstract
Critics are applied to estimating policy gradients in the actor-critic framework, an essential part of Reinforcement Learning methods. An appropriate critic is supposed to balance variance from sample returns and bias introduced by parameterized value functions. A typical critic of balancing variance and bias, the generalized advantage estimator (GAE), combines sample returns and value functions with a fixed weight parameter. However, such a parameter is hard to fit, which results in no promised stability for GAE. In this paper, indicators of variance and bias are proposed to get adaptive weight parameters, with which adaptive advantage estimators are obtained. Empirical results on both 2D and 3D simulated robotic locomotion tasks show that the adaptive advantage estimators achieve similar or superior performance compared to GAE.
Yurou Chen, Zhiyong Liu 0001
IJCNN3
2021 Zero-shot policy generation in lifelong reinforcement learning
Yiming Qian, Fangzhou Xiong, Zhiyong Liu 0001
Neurocomputing3
2021 Graph matching based point correspondence with alternating direction method of multipliers
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001, Mingyu Fan
Neurocomputing4
2021 Supervised learning for parameterized Koopmans-Beckmann's graph matching
Shaofeng Zeng, Zhiyong Liu 0001, Xu Yang 0004
Pattern Recognit. Lett.2
2021 A Doubly Graduated Method for Inference in Markov Random Field
abstract
Maximum a posteriori (MAP) inference in Markov random field (MRF) lays the foundation for many computer vision tasks, which can be formulated by a binary quadratic programming (BQP) problem. Compared with the discrete methods, the continuous relaxation scheme becomes popular due to its generality and efficiency. However, existing continuous relaxation based MAP algorithms are still limited by two problems, i.e., the highly nonconvex original objective function and the gap between the original BQP problem and the relaxed continuous optimization problem. Targeting the two problems, this paper presents a doubly graduated continuous relaxation algorithm for MAP inference in MRF, which are, respectively, the Gaussian smoothing based graduated nonconvexity process and conditional gradient ascent based graduated projection. Experiments on both synthetic data and real-world images illustrate the algorithm's state-of-the-art performance in objective function optimization and typical computer vision tasks.
Xu Yang 0004, Zhiyong Liu 0001
SIAM J. Imaging Sci.2
2020 Intra-domain Knowledge Generalization in Cross-Domain Lifelong Reinforcement Learning
Yiming Qian, Fangzhou Xiong, Zhiyong Liu 0001
ICONIP (5)3
2020 Knowledge-Experience Graph with Denoising Autoencoder for Zero-Shot Learning in Visual Cognitive Development
Xu Yang 0004, Zhiyong Liu 0001, Lu Zhang 0054, Dongchun Ren, Mingyu Fan
ICONIP (5)3
2020 MixedFusion: 6D Object Pose Estimation from Decoupled RGB-Depth Features
abstract
Estimating the 6D pose of objects is an important process for intelligent systems to achieve interaction with the real-world. As the RGB-D sensors become more accessible, the fusion-based methods have prevailed, since the point clouds provide complementary geometric information with RGB values. However, due to the difference in feature space between color image and depth image, the network structures that directly perform point-to-point matching fusion do not effectively fuse the features of the two. In this paper, we propose a simple but effective approach, named MixedFusion. Different from the prior works, we argue that the spatial correspondence of color and point clouds could be decoupled and reconnected, thus enabling a more flexible fusion scheme. By performing the proposed method, more informative points can be mixed and fused with rich color features. Extensive experiments are conducted on the challenging LineMod and YCB-Video datasets, which shows that our method significantly boosts the performance without introducing extra overheads. Furthermore, when the minimum tolerance of metric narrows, the proposed approach performs better for the high-precision demands.
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001
ICPR4
2020 Encoding primitives generation policy learning for robotic arm to overcome catastrophic forgetting in sequential multi-tasks learning
Fangzhou Xiong, Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Hong Qiao, Amir Hussain 0001
Neural Networks2
2020 A Continuation Method for Graph Matching Based Feature Correspondence
abstract
Feature correspondence lays the foundation for many computer vision and image processing tasks, which can be well formulated and solved by graph matching. Because of the high complexity, approximate methods are necessary for graph matching, and the continuous relaxation provides an efficient approximate scheme. But there are still many problems to be settled, such as the highly nonconvex objective function, the ignorance of the combinatorial nature of graph matching in the optimization process, and few attention to the outlier problem. Focusing on these problems, this paper introduces a continuation method directly targeting at the combinatorial optimization problem associated with graph matching. Specifically, first a regularization function incorporating the original objective function and the discrete constraints is proposed. Then a continuation method based on Gaussian smoothing is applied to it, in which the closed forms of relevant functions with respect to the outlier distribution are deduced. Experiments on both synthetic data and real world images validate the effectiveness of the proposed method.
Xu Yang 0004, Zhiyong Liu 0001, Hong Qiao
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Incorporating Discrete Constraints Into Random Walk-Based Graph Matching
abstract
Graph matching is a fundamental problem in theoretical computer science and artificial intelligence, and lays the foundation for many computer vision and machine learning tasks. Approximate algorithms are necessary for graph matching due to its NP-complete nature. Inspired by the usage in network-related tasks, random walk is generalized to graph matching as a type of approximate algorithm. However, it may be inappropriate for the previous random walk-based graph matching algorithms to utilize continuous techniques without considering the discrete property. In this paper, we propose a novel random walk-based graph matching algorithm by incorporating both continuous and discrete constraints in the optimization process. Specifically, after interpreting graph matching by random walk, the continuous constraints are directly embedded in the random walk constraint in each iteration. Further, both the assignment matrix (vector) and the pairwise similarity measure between graphs are iteratively updated according the discrete constraints, which automatically leads the continuous solution to the discrete domain. Comparisons on both synthetic and real-world data demonstrate the effectiveness of the proposed algorithm.
Xu Yang 0004, Zhiyong Liu 0001, Hong Qiaoxu
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not strictly aligned, making one object has different positions in different modalities. In deep learning based methods, this problem makes it difficult to fuse the feature maps from both modalities and puzzles the CNN training. In this paper, we propose a novel Aligned Region CNN (AR-CNN) to handle the weakly aligned multispectral data in an end-to-end way. Firstly, we design a Region Feature Alignment (RFA) module to capture the position shift and adaptively align the region features of the two modalities. Secondly, we present a new multimodal fusion method, which performs feature re-weighting to select more reliable features and suppress the useless ones. Besides, we propose a novel RoI jitter strategy to improve the robustness to unexpected shift patterns of different devices and system settings. Finally, since our method depends on a new kind of labelling: bounding boxes that match each modality, we manually relabel the KAIST dataset by locating bounding boxes in both modalities and building their relationships, providing a new KAIST-Paired Annotation. Extensive experimental validations on existing datasets are performed, demonstrating the effectiveness and robustness of the proposed method. Code and data are available at https://github.com/luzhang16/AR-CNN.
Lu Zhang 0054, Xiangyu Zhu 0001, Xu Yang 0004, Zhen Lei 0001, Zhiyong Liu 0001
ICCV6
2019 Learning Transferable Policies with Improved Graph Neural Networks on Serial Robotic Structure
Fangzhou Xiong, Xu Yang 0004, Zhiyong Liu 0001
ICONIP (3)4
2019 Sub-hypergraph matching based on adjacency tensor
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001
Comput. Vis. Image Underst.4
2019 Special issue on advances in graph algorithm and applications
Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Cheng-Lin Liu 0001
Neurocomputing1
2019 Caging a novel object using multi-task learning method
Jianhua Su, Hong Qiao, Zhiyong Liu 0001
Neurocomputing4
2019 Special issue on deep learning for intelligent sensing, decision-making and control
Wei Zhang 0021, Junchi Yan, Zhiyong Liu 0001, Zhigang Zeng
Neurocomputing3
2019 Guided Policy Search for Sequential Multitask Learning
abstract
Policy search in reinforcement learning (RL) is a practical approach to interact directly with environments in parameter spaces, that often deal with dilemmas of local optima and real-time sample collection. A promising algorithm, known as guided policy search (GPS), is capable of handling the challenge of training samples using trajectory-centric methods. It can also provide asymptotic local convergence guarantees. However, in its current form, the GPS algorithm cannot operate in sequential multitask learning scenarios. This is due to its batch-style training requirement, where all training samples are collectively provided at the start of the learning process. The algorithm’s adaptation is thus hindered for real-time applications, where training samples or tasks can arrive randomly. In this paper, the GPS approach is reformulated, by adapting a recently proposed, lifelong-learning method, and elastic weight consolidation. Specifically, Fisher information is incorporated to impart knowledge from previously learned tasks. The proposed algorithm, termed sequential multitask learning-GPS, is able to operate in sequential multitask learning settings and ensuring continuous policy learning, without catastrophic forgetting. Pendulum and robotic manipulation experiments demonstrate the new algorithms efficacy to learn control policies for handling sequentially arriving training samples, delivering comparable performance to the traditional, and batch-based GPS algorithm. In conclusion, the proposed algorithm is posited as a new benchmark for the real-time RL and robotics research community.
Fangzhou Xiong, Biao Sun 0005, Xu Yang 0004, Hong Qiao, Kaizhu Huang, Amir Hussain 0001, Zhiyong Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.7
2018 Overcoming Catastrophic Forgetting with Self-adaptive Identifiers
Fangzhou Xiong, Zhiyong Liu 0001, Xu Yang 0004
ICONIP (3)2
2018 Graph Matching Based on Fast Normalized Cut
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001
ICONIP (6)4
2018 Single Shot Feature Aggregation Network for Underwater Object Detection
abstract
The rapidly developing ocean exploration and observation make the demand for underwater object detection become increasingly urgent. Recently, deep convolutional neural networks (CNN) have shown strong ability in feature representation and CNN-based detectors also achieve remarkable performance, but still facing the big challenge when detecting multi-scale objects in a complex underwater environment. To address this challenge, we propose a novel underwater object detector, introducing multiscale features and complementary context information for better classification and location ability. In the auto-grabbing contest of 2017 Underwater Robot Picking Contest sponsored by National Natural Science Foundation of China (NSFC), we won the 1-st place by using proposed method for real coastal underwater object detection.
Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001, Lu Qi 0001, Hao Zhou 0014, Charles Chiu
ICPR3
2018 Adaptive Graph Matching
abstract
Establishing correspondence between point sets lays the foundation for many computer vision and pattern recognition tasks. It can be well defined and solved by graph matching. However, outliers may significantly deteriorate its performance, especially when outliers exist in both point sets and meanwhile the inlier number is unknown. In this paper, we propose an adaptive graph matching algorithm to tackle this problem. Specifically, a novel formulation is proposed to make the graph matching model adaptively determine the number of inliers and match them, then by relaxing the discrete domain to its convex hull the discrete optimization problem is relaxed to be a continuous one, and finally a graduated projection scheme is used to get a discrete matching solution. Consequently, the proposed algorithm could realize inlier number estimation, inlier selection, and inlier matching in one optimization framework. Experiments on both synthetic data and real world images witness the effectiveness of the proposed algorithm.
Xu Yang 0004, Zhiyong Liu 0001
IEEE Trans. Cybern.2
2018 An Algorithm for Finding the Most Similar Given Sized Subgraphs in Two Weighted Graphs
abstract
We propose a weighted common subgraph (WCS) matching algorithm to find the most similar subgraphs in two labeled weighted graphs. WCS matching, as a natural generalization of equal-sized graph matching and subgraph matching, has found wide applications in many computer vision and machine learning tasks. In this brief, WCS matching is first formulated as a combinatorial optimization problem over the set of partial permutation matrices. Then, it is approximately solved by a recently proposed combinatorial optimization framework-graduated nonconvexity and concavity procedure. Experimental comparisons on both synthetic graphs and real-world images validate its robustness against noise level, problem size, outlier number, and edge density.
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2017 A Linear Online Guided Policy Search Algorithm
Biao Sun 0005, Fangzhou Xiong, Zhiyong Liu 0001, Xu Yang 0004, Hong Qiao
ICONIP (5)3
2017 A Bayesian Posterior Updating Algorithm in Reinforcement Learning
Fangzhou Xiong, Zhiyong Liu 0001, Xu Yang 0004, Biao Sun 0005, Charles Chiu, Hong Qiao
ICONIP (5)2
2017 LibCoopt: A library for combinatorial optimization on partial permutation matrices
Zhiyong Liu 0001, Xu Yang 0004, Shi-Hao Feng
Neurocomputing1
2017 Probabilistic hypergraph matching based on affinity tensor updating
Xu Yang 0004, Zhiyong Liu 0001, Hong Qiao, Jianhua Su
Neurocomputing2
2017 Point correspondence by a new third order graph matching algorithm
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001
Pattern Recognit.3
2016 A Novel Manifold Regularized Online Semi-supervised Learning Algorithm
Shuguang Ding, Xuanyang Xi, Zhiyong Liu 0001, Hong Qiao, Bo Zhang 0006
ICONIP (1)3
2016 Stitching contaminated images
Chuan Li 0004, Zhiyong Liu 0001, Xu Yang 0004, Hong Qiao, Jianhua Su
Neurocomputing2
2016 Introducing locally affine-invariance constraints into lunar surface image correspondence
Yuren Zhang, Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001, Chuankai Liu
Neurocomputing4
2016 Large Scale Online Kernel Learning
abstract
In this paper, we present a new framework for large scale online kernel learning, making kernel methods efficient and scalable for large-scale online learning applications. Unlike the regular budget online kernel learning scheme that usually uses some budget maintenance strategies to bound the number of support vectors, our framework explores a completely different approach of kernel functional approximation techniques to make the subsequent online learning task efficient and scalable. Specifically, we present two different online kernel machine learning algorithms: (i) Fourier Online Gradient Descent (FOGD) algorithm that applies the random Fourier features for approximating kernel functions; and (ii) Nyström Online Gradient Descent (NOGD) algorithm that applies the Nyström method to approximate large kernel matrices. We explore these two approaches to tackle three online learning tasks: binary classification, multi-class classification, and regression. The encouraging results of our experiments on large-scale datasets validate the effectiveness and efficiency of the proposed algorithms, making them potentially more practical than the family of existing budget online kernel learning approaches.
Steven C. H. Hoi, Peilin Zhao, Zhiyong Liu 0001
J. Mach. Learn. Res.5
2016 Online Multi-Modal Distance Metric Learning with Application to Image Retrieval
abstract
Distance metric learning (DML) is an important technique to improve similarity search in content-based image retrieval. Despite being studied extensively, most existing DML approaches typically adopt a single-modal learning framework that learns the distance metric on either a single feature type or a combined feature space where multiple types of features are simply concatenated. Such single-modal DML methods suffer from some critical limitations: (i) some type of features may significantly dominate the others in the DML task due to diverse feature representations; and (ii) learning a distance metric on the combined high-dimensional feature space can be extremely time-consuming using the naive feature concatenation approach. To address these limitations, in this paper, we investigate a novel scheme of online multi-modal distance metric learning (OMDML), which explores a unified two-level online learning scheme: (i) it learns to optimize a distance metric on each individual feature space; and (ii) then it learns to find the optimal combination of diverse types of features. To further reduce the expensive cost of DML on high-dimensional feature space, we propose a low-rank OMDML algorithm which not only significantly reduces the computational cost but also retains highly competing or even better learning accuracy. We conduct extensive experiments to evaluate the performance of the proposed algorithms for multi-modal image retrieval, in which encouraging results validate the effectiveness of the proposed technique.
Steven C. H. Hoi, Peilin Zhao, Chunyan Miao, Zhiyong Liu 0001
IEEE Trans. Knowl. Data Eng.5
2015 Moving average reversion strategy for on-line portfolio selection
abstract
On-line portfolio selection, a fundamental problem in computational finance, has attracted increasing interest from artificial intelligence and machine learning communities in recent years. Empirical evidence shows that stock's high and low prices are temporary and stock prices are likely to follow the mean reversion phenomenon. While existing mean reversion strategies are shown to achieve good empirical performance on many real datasets, they often make the single-period mean reversion assumption, which is not always satisfied, leading to poor performance in certain real datasets. To overcome this limitation, this article proposes a multiple-period mean reversion, or so-called “Moving Average Reversion” (MAR), and a new on-line portfolio selection strategy named “On-Line Moving Average Reversion” (OLMAR), which exploits MAR via efficient and scalable online machine learning techniques. From our empirical results on real markets, we found that OLMAR can overcome the drawbacks of existing mean reversion algorithms and achieve significantly better results, especially on the datasets where existing mean reversion algorithms failed. In addition to its superior empirical performance, OLMAR also runs extremely fast, further supporting its practical applicability to a wide range of applications. Finally, we have made all the datasets and source codes of this work publicly available at our project website: http://OLPS.stevenhoi.org/.
Bin Li 0027, Steven C. H. Hoi, Doyen Sahoo, Zhiyong Liu 0001
Artif. Intell.4
2015 Feature correspondence based on directed structural model matching
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001
Image Vis. Comput.3
2015 Outlier robust point correspondence based on GNCCP
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001
Pattern Recognit. Lett.3
2015 Vision-Based Caging Grasps of Polyhedron-Like Workpieces With a Binary Industrial Gripper
abstract
Development of a flexible low-cost robotic system to grasp various 3D workpieces is of practical important. Traditional approaches on 3D grasping usually need to compute the effective contact positions based on sufficient conditions, such as “force-closure” or “form-closure,” to prevent all the motions of the grasped objects. Compared with the previous work motivated by high-precision applications, caging provides a way to manipulate an object without needing to immobilize it. However, most caging conditions consider only 2D motions of the object in grasping. In some cases, other motions in 3D space, e.g., the pitch and roll rotations of the objects, would also possibly change a caging configuration to an uncaging configuration. This paper aims to discuss caging with frictionless contact, by taking into consideration of all the motions of the object. We first discuss the relationship between the gripper configuration and the state of the grasped object in grasping, where all the motions of the object are taken into account. We then establish the sufficient conditions to find a set of 3D caging configurations in terms of the width of the object's projection and the gap of the gripper, such that we can utilize the projection to find the feasible placements of the pins to determine the caging configuration. Furthermore, we discuss how to find the caging configuration that is capable of leading to a form-closure based on attractive region formed in the configuration space.
Jianhua Su, Hong Qiao, Zhicai Ou, Zhiyong Liu 0001
IEEE Trans Autom. Sci. Eng.4
2014 MAP Inference with MRF by Graduated Non-Convexity and Concavity Procedure
Zhiyong Liu 0001, Hong Qiao, Jianhua Su
ICONIP (2)1
2014 Graph Matching by Simplified Convex-Concave Relaxation Procedure
Zhiyong Liu 0001, Hong Qiao, Xu Yang 0004, Steven C. H. Hoi
Int. J. Comput. Vis.1
2014 A graph matching algorithm based on concavely regularized convex relaxation
Zhiyong Liu 0001, Hong Qiao, Li-Hao Jia, Lei Xu 0001
Neurocomputing1
2014 GNCCP - Graduated NonConvexityand Concavity Procedure
abstract
In this paper we propose the graduated nonconvexity and concavity procedure (GNCCP) as a general optimization framework to approximately solve the combinatorial optimization problems defined on the set of partial permutation matrices. GNCCP comprises two sub-procedures, graduated nonconvexity which realizes a convex relaxation and graduated concavity which realizes a concave relaxation. It is proved that GNCCP realizes exactly a type of convex-concave relaxation procedure (CCRP), but with a much simpler formulation without needing convex or concave relaxation in an explicit way. Actually, GNCCP involves only the gradient of the objective function and is therefore very easy to use in practical applications. Two typical related NP-hard problems, partial graph matching and quadratic assignment problem (QAP), are employed to demonstrate its simplicity and state-of-the-art performance.
Zhiyong Liu 0001, Hong Qiao
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Large Scale Online Kernel Classification
Steven C. H. Hoi, Peilin Zhao, Jinfeng Zhuang, Zhiyong Liu 0001
IJCAI5
2013 Online multi-task collaborative filtering for on-the-fly recommender systems
abstract
Traditional batch model-based Collaborative Filtering (CF) approaches typically assume a collection of users' rating data is given a priori for training the model. They suffer from a common yet critical drawback, i.e., the model has to be re-trained completely from scratch whenever new training data arrives, which is clearly non-scalable for large real recommender systems where users' rating data often arrives sequentially and frequently. In this paper, we investigate a novel efficient and scalable online collaborative filtering technique for on-the-fly recommender systems, which is able to effectively online update the recommendation model from a sequence of rating observations. Specifically, we propose a family of online multi-task collaborative filtering (OMTCF) algorithms, which tackle the online collaborative filtering task by exploiting the similar principle as online multitask learning. Encouraging empirical results on large-scale datasets showed that the proposed technique is significantly more effective than the state-of-the-art algorithms.
Steven C. H. Hoi, Peilin Zhao, Zhiyong Liu 0001
RecSys4
2013 Partial correspondence based on subgraph matching
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001
Neurocomputing3
2012 Novel DR-tree index based on the diagonal line of MBR
abstract
The application of spatial database is increasingly widespread. How to effectively store and organize multidimensional space data and improve the processing efficiency of multidimensional data has become a central issue. R-tree is one of the most widely used spatial indexes. Due to the existing large overlap and coverage among the nodes of R-tree, the search path of a data object is not unique, and the search efficiency declines sharply when the amount of data increases. Based on the analysis and research of previous index trees, a novel DR-tree index based on the diagonal lines of MBR is proposed in this paper. DR-tree uses the diagonal line of MBR to indicate spatial data objects and construct R-tree, still adopt the endpoint coordinates of MBR diagonal to signify the location of leaf nodes or non-leaf nodes. Since MBR is simplified, the coverage and overlap among regions are also reduced. Experimental results show that the new index tree is superior to R-tree in performance. The query paths of data objects are single, the insertion, deletion, and query efficiency of data objects are significantly improved, and the performance of the new index tree becomes more apparent when the amount of data objects increases.
Jine Tang, Zhangbing Zhou, Zhiyong Liu 0001
IWCMC3
2012 An Extended Path Following Algorithm for Graph-Matching Problem
abstract
The path following algorithm was proposed recently to approximately solve the matching problems on undirected graph models and exhibited a state-of-the-art performance on matching accuracy. In this paper, we extend the path following algorithm to the matching problems on directed graph models by proposing a concave relaxation for the problem. Based on the concave and convex relaxations, a series of objective functions are constructed, and the Frank-Wolfe algorithm is then utilized to minimize them. Several experiments on synthetic and real data witness the validity of the extended path following algorithm.
Zhiyong Liu 0001, Hong Qiao, Lei Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Investigation on the skewness for independent component analysis
Zhiyong Liu 0001, Hong Qiao
Sci. China Inf. Sci.1
2009 Multiple ellipses detection in noisy environments: A hierarchical approach
Zhiyong Liu 0001, Hong Qiao
Pattern Recognit.1
2007 A Robust Multiple Cues Fusion based Bayesian Tracker
abstract
This paper presents an efficient and robust tracking algorithm based on multiple cues fusion in the Bayesian framework. This method characterizes the object to be tracked using a MOG (mixture of Gaussians) based appearance model and a chamfer-matching based shape model. A selective updating technique for the models is employed to accommodate for appearance and illumination changes. Meantime, the mean shift algorithm is embedded as the prior information into the Bayesian framework to give a heuristic prediction in the hypotheses generation process, which also alleviates the great computational load suffered by the conventional Bayesian tracker. Experimental results demonstrate that, compared with some existing works, the proposed algorithm has a better adaptability to changes of the object as well as the environments.
Xiaoqin Zhang 0002, Zhiyong Liu 0001, Hong Qiao
ICRA2
2007 Investigation on Multisets Mixture Learning Based Object Detection
abstract
By minimizing the mean square reconstruction error, multisets mixture learning (MML) provides a general approach for object detection in image. To calculate each sample reconstruction error, as the object template is represented by a set of contour points, the MML needs to inefficiently enumerate the distances between the sample and all the contour points. In this paper, we develop the line segment approximation (LSA) algorithm to calculate the reconstruction error, which is shown theoretically and experimentally to be more efficient than the enumeration method. It is also experimentally illustrated that the MML based algorithm has a better noise resistance ability than the generalized Hough transform (GHT) based counterpart.
Zhiyong Liu 0001, Hong Qiao, Lei Xu 0001
Int. J. Pattern Recognit. Artif. Intell.1
2006 Multi-Information Fusion for Scale Selection in Robot Tracking
abstract
Mean shift, for its simplicity and efficiency, has achieved a considerable success in robot tracking. For the mean shift based tracking algorithm, the scale of the mean-shift kernel bandwidth is a crucial parameter which reflects the size of tracking window. However, in literature how to properly update or select the bandwidth remains a tough task as the size of the object under consideration changes. In this paper, a weighted average integral projection approach is proposed to extract the local information of the object, and then a multiinformation fusion strategy is suggested for the scale selection, which combines both the global and local information of the sample weight image. Moreover, a coarse-to-fine approximate approach is employed to accelerate the procedure. Experimental results demonstrate that, compared to some existing works, the strategy proposed has a better adaptability as the size of the object changes in clutter environments.
Xiaoqin Zhang 0002, Hong Qiao, Zhiyong Liu 0001
IROS3
2006 Multisets mixture learning-based ellipse detection
Zhiyong Liu 0001, Hong Qiao, Lei Xu 0001
Pattern Recognit.1