EDBT 2026 Demo / reviewers in the wild / expert
Xu Yang 0004
dblp:63/1534-4
· DBLP profile ↗
63ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0003-0553-4581ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 9 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Multimodal Alignment for Image-Text Retrieval via Graph Learning
Mao Chen 0007, Xiangkai Zhang, Lu Qi 0001, Xiangtai Li, Xu Yang 0004, Steven C. H. Hoi, Zhiyong Liu 0001, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | Collaborative Observation Imputation and Trajectory Prediction via Consistency EvaluationabstractPedestrian trajectory prediction is a critical task in various mobile computing applications, such as video surveillance, robot navigation, autonomous driving, and human mobility analysis. Although significant progress has been made by current methods, the challenge of observation deficiency in pedestrian trajectory prediction remains largely unaddressed. Since most existing methods focus on optimizing prediction accuracy under the assumption of complete observations, while ignoring the potential for observation deficiency caused by failures in detection or tracking algorithms. To overcome this challenge, we propose a collaborative observation imputation and trajectory prediction framework, which employs consistency evaluation to jointly perform the imputation and prediction tasks. Specifically, we first build a consistency evaluation module to align features between observed and future trajectory pairs using contrastive learning. Then, we design a trajectory imputation and prediction baseline, which adopts a parallel paradigm, to mitigate the impact of coarse imputations on trajectory prediction when performing initial imputation and prediction. Next, we introduce a consistency-guided Skip-Diffusion module, which leverages consistency evaluation between initial imputations and ground truth future trajectories to refine the initial imputations. Finally, we propose a consistency-driven Cross-Mamba module, which uses consistency evaluation between ground-truth observations and initial predictions to refine the initial predictions. Extensive experiments demonstrate the effectiveness of the proposed framework in both imputation and prediction tasks. Hao Zhou 0014, Mingyu Fan, Xu Yang 0004, Hai Huang 0004, Kaishun Wu, Lu Wang 0002, Fei Luo 0003 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Mimic In-Context Learning for Multimodal TasksabstractRecently, In-context Learning (ICL) has become a significant inference paradigm in Large Multimodal Models (LMMs), utilizing a few in-context demonstrations (ICDs) to prompt LMMs for new tasks. However, the synergistic effects in multimodal data increase the sensitivity of ICL performance to the configurations of ICDs, stimulating the need for a more stable and general mapping function. Mathematically, in Transformer-based models, ICDs act as "shift vectors" added to the hidden states of query tokens. Inspired by this, we introduce Mimic In-Context Learning (MimIC) to learn stable and generalizable shift effects from ICDs. Specifically, compared with some previous shift vector-based methods, MimIC more strictly approximates the shift effects by integrating lightweight learnable modules into LMMs with four key enhancements: 1) inserting shift vectors after attention layers, 2) assigning a shift vector to each attention head, 3) making shift magnitude query-dependent, and 4) employing a layer-wise alignment loss. Extensive experiments on two LMMs (Idefics-9b and Idefics2-8b-base) across three multimodal tasks (VQAv2, OK-VQA, Captioning) demonstrate that MimIC outperforms existing shift vector-based methods. The code is available at https://github.com/Kamichanw/MimIC. Yuchu Jiang, Jiale Fu, Chenduo Hao, Xinting Hu, Yingzhe Peng, Xin Geng 0001, Xu Yang 0004 |
CVPR | 7 |
| 2025 | Number it: Temporal Grounding Videos like Flipping MangaabstractVideo Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue. However, they struggle to extend this visual understanding to tasks requiring precise temporal localization, known as Video Temporal Grounding (VTG). To address this, we introduce Number-Prompt (NumPro), a novel method that empowers Vid-LLMs to bridge visual comprehension with temporal grounding by adding unique numerical identifiers to each video frame. Treating a video as a sequence of numbered frame images, NumPro transforms VTG into an intuitive process: flipping through manga panels in sequence. This allows Vid-LLMs to “read” event timelines, accurately linking visual content with cor responding temporal information. Our experiments demonstrate that NumPro significantly boosts VTG performance of top-tier Vid-LLMs without additional computational cost. Furthermore, fine-tuning on a NumPro-enhanced dataset defines a new state-of-the-art for VTG, surpassing previous top-performing methods by up to 6.9% in mIoU for moment retrieval and 8.5% in mAP for highlight detection. The code is available at https://github.com/yongliang-wu/NumPro. Yongliang Wu, Xinting Hu, Yizhou Zhou, Fengyun Rao, Bernt Schiele, Xu Yang 0004 |
CVPR | 8 |
| 2025 | Language-Conditioned Waypoint Predictor for Continuous Vision-and-Language NavigationabstractWaypoint prediction is a popular technique for Vision-and-Language Navigation in Continuous Environments (VLN-CE), which abstracts navigable locations as waypoints to ease the subsequent action prediction. Nevertheless, we found current waypoint predictors are not always accurate, limiting navigation’s overall performance. One possible reason may be the lack of language context, leading to the failure to generate corresponding waypoints for critical locations mentioned in the instructions. To that end, we propose a novel framework to enable the training of the language-conditioned waypoint predictor. First, as the VLN-CE agents ground instructions with the environment when navigating, we employ a pre-trained agent to encode language for the waypoint predictor. Second, the language-conditioned waypoint predictor is trained with the data collected using the same agent. Third, we train the new VLN-CE navigation agent with the proposed waypoint predictor. Fourth, the disparity between the language encoder agent and the navigation agent drives us to devise a cycle training scheme to alternately train the agent and the waypoint predictor, further enhancing the performance of both the waypoint predictor and navigation agent. Experimental results show that our waypoint predictor’s performance surpasses all existing ones. With better waypoints, the gap between waypoint-based methods and their upper bound narrows by about 60%. Yuankai Qi, Xu Yang 0004, Zhaoxiang Zhang 0001 |
ICME | 4 |
| 2025 | Three-Dimensional Trajectory Prediction with 3DMoTraj DatasetabstractWith the growing interest in embodied and spatial intelligence, accurately predicting trajectories in 3D environments has become increasingly critical. However, no datasets have been explicitly designed to study 3D trajectory prediction. To this end, we contribute a 3D motion trajectory (3DMoTraj) dataset collected from unmanned underwater vehicles (UUVs) operating in oceanic environments. Mathematically, trajectory prediction becomes significantly more complex when transitioning from 2D to 3D. To tackle this challenge, we analyze the prediction complexity of 3D trajectories and propose a new method consisting of two key components: decoupled trajectory prediction and correlated trajectory refinement. The former decouples inter-axis correlations, thereby reducing prediction complexity and generating coarse predictions. The latter refines the coarse predictions by modeling their inter-axis correlations. Extensive experiments show that our method significantly improves 3D trajectory prediction accuracy and outperforms state-of-the-art methods. Both the 3DMoTraj dataset and the method are available at https://github.com/zhouhao94/3DMoTraj. Hao Zhou 0014, Xu Yang 0004, Mingyu Fan, Lu Qi 0001, Xiangtai Li, Ming-Hsuan Yang 0001, Fei Luo 0003 |
ICML | 2 |
| 2025 | CARD: Control-Driven Autoregressive Reconstruction with Decoupled Learning for Multi-Class Anomaly DetectionabstractMulti-class unsupervised anomaly detection (UAD) is challenging due to the difficulty of harmonizing distributional differences across categories within a unified framework. While recent diffusion-based methods have demonstrated promising performance by leveraging denoising processes for anomaly reconstruction, the lack of explicit causal constraints limits their ability to handle complex logical inconsistencies. Inspired by the success of autoregressive models in enforcing local-to-global logical consistency, we propose a Control-driven Autoregressive Reconstruction with Decoupled learning (CARD) framework for multi-class UAD. It first tokenizes images using vector quantization and employs a vision autoregressive model to capture causal dependencies within normal patterns as explicit prior knowledge. Then, we introduce a Control-Driven Reconstruction (CDR) network, which aligns input features as control signals into the frozen autoregressive model to adjust the predicted distribution, enabling a generative reconstruction process. Additionally, we apply perturbations to the CDR input to simulate anomalous conditions, facilitating the model to correct out-of-distribution anomaly features under prior knowledge constraints. By decoupling normal pattern learning from reconstruction, CARD prevents identity mapping caused by forgetting implicit priors in the conventional reconstruction-based method. Comprehensive experiments on several benchmark datasets validate the effectiveness of our approach. On the MVTecAD dataset, CARD achieved AUROC scores of 98.7% at the image level and 98.5% at the pixel level. Mingqing Wang, Boyi Sun, Qianfan Zhao, Lu Zhang 0054, Zhiyong Liu 0001, Xu Yang 0004, Suiwu Zheng |
MMAsia | 7 |
| 2025 | RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Embodied IntelligenceabstractThis paper addresses the scarcity of low-cost but high-dexterity platforms for collecting real-world multi-fingered robot manipulation data towards generalist robot autonomy.
To achieve it, we propose the RAPID Hand, a co-optimized hardware and software platform where the compact 20-DoF hand, robust whole-hand perception, and high-DoF teleoperation interface are jointly designed.
Specifically, RAPID Hand adopts a compact and practical hand ontology and a hardware-level perception framework that stably integrates wrist-mounted vision, fingertip tactile sensing, and proprioception with sub-7 ms latency and spatial alignment.
Collecting high-quality demonstrations on high-DoF hands is challenging, as existing teleoperation methods struggle with precision and stability on complex multi-fingered systems.
We address this by co-optimizing hand design, perception integration, and teleoperation interface through a universal actuation scheme, custom perception electronics, and two retargeting constraints. We evaluate the platform’s hardware, perception, and teleoperation interface. Training a diffusion policy on collected data shows superior performance over prior works, validating the system’s capability for reliable, high-quality data collection.
The platform is constructed from low-cost and off-the-shelf components and will be made public to ensure reproducibility and ease of adoption. Zhaoliang Wan, Zetong Bi, Zida Zhou, Hao Ren 0006, Yiming Zeng 0008, Lu Qi 0001, Xu Yang 0004, Ming-Hsuan Yang 0001, Hui Cheng 0002 |
NeurIPS | 8 |
| 2025 | KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing ModelsabstractRecent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, We introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on nine state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems. Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye, Xianfang Zeng, Bernt Schiele, Ming-Hsuan Yang 0001, Xu Yang 0004 |
NeurIPS | 10 |
| 2025 | Rethinking Evaluation Metrics of Open-Vocabulary SegmentationabstractThis paper highlights a problem of evaluation metrics adopted in the open-vocabulary segmentation. The evaluation process relies heavily on closed-set metrics on zero-shot or cross-dataset pipelines without considering the similarity between predicted and ground truth categories. We first survey eleven similarity measurements between two categorical words using WordNet linguistics statistics, text embedding, or language models by comprehensive quantitative analysis and user study to tackle this issue. Based on those explored measurements, we design novel evaluation metrics, Open mIoU, Open AP, and Open PQ, tailored for three open-vocabulary segmentation tasks. We benchmark the proposed evaluation metrics on twelve open-vocabulary methods in three segmentation tasks. Despite the relative subjectivity of similarity distance, we demonstrate that our metrics can still well evaluate the open ability of the existing open-vocabulary segmentation methods. We hope our work can bring the community new thinking about evaluating model ability for open-vocabulary segmentation. Hao Zhou 0014, Lu Qi 0001, Tiancheng Shen, Hai Huang 0004, Xu Yang 0004, Xiangtai Li, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | DeepPartitioning: Deep Learning of Graph Partitioning for Neuron Segmentation From Electron Microscopy Volume via Graph Neural NetworkabstractSuperpixel aggregation represents a highly effective approach for automated neuron segmentation from electron microscopy (EM) volumes, which can be considered as a graph partitioning task on the region adjacency graph (RAG) of extracted superpixels. However, existing graph partitioning models for superpixel aggregation suffer from the modeling error due to insufficient model capacity. More specifically, the modeling error is caused by the simplification in formulating the real-world graph partitioning task (i.e., superpixel aggregation) into a mathematically well-defined optimization problem. To address this issue, we sidestep the explicit formulation and propose a fully end-to-end superpixel aggregation method based on deep learning of the graph partitioning task, called DeepPartitioning. The central challenge lies in characterizing the partitioning task involving combinatorial complexity. Hence, our method incorporates a line graph neural network (LGNN) to capture higher-order relational structures in RAGs. Specifically, the LGNN enables the propagation of second-order superpixel-pair features among adjacent edges in RAGs. In this way, the partitioning task can be implicitly transformed into the vanilla second-order multicut problem while maintaining higher-order structural information. Overall, our method integrates a second-order feature extractor, a higher-order feature integrator (i.e., the LGNN), and a differentiable approximation to a multicut solver into a unified, learnable framework. Extensive experiments on three public EM datasets demonstrate the effectiveness of the proposed DeepPartitioning within the neuron segmentation pipeline. Zhenchen Li, Xu Yang 0004, Jing Liu 0054, Zhiyong Liu 0001, Hua Han 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Deep Graph Reinforcement Learning for Solving Multicut ProblemabstractThe multicut problem, also known as correlation clustering, is a classic combinatorial optimization problem that aims to optimize graph partitioning given only node (dis)similarities on edges. It serves as an elegant generalization for several graph partitioning problems and has found successful applications in various areas such as data mining and computer vision. However, the multicut problem with an exponentially large number of cycle constraints proves to be NP-hard, and existing solvers either suffer from exponential complexity or often give unsatisfactory solutions due to inflexible heuristics driven by hand-designed mechanisms. In this article, we propose a deep graph reinforcement learning method to solve the multicut problem within a combinatorial decision framework involving sequential edge contractions. The customized subgraph neural network adapts to the dynamically edge-contracted graph environment by extracting bilevel connected features from both contracted and original graphs. Our method can learn to infer feasible multicut solutions end-to-end toward optimization of the multicut objective in a data-driven manner. More specifically, by exploring the decision space adaptively, it implicitly gains heuristic knowledge from topological patterns of instances and thereby generates more targeted heuristics overcoming the short-sightedness inherent in the hand-designed ones. During testing, the learned heuristics iteratively contract graphs to construct high-quality solutions within polynomial time. Extensive experiments on synthetic and real-world multicut instances show the superiority of our method over existing combinatorial solvers, while also maintaining a certain level of out-of-distribution generalization ability. Zhenchen Li, Xu Yang 0004, Shaofeng Zeng, Jingbin Yuan, Zhiyong Liu 0001, Hua Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Weakly Aligned Feature Fusion for Multimodal Object DetectionabstractTo achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image pair is not strictly aligned, making one object has different positions in different modalities. For the deep learning method, this problem makes it difficult to fuse multimodal features and puzzles the convolutional neural network (CNN) training. In this article, we propose a general multimodal detector named aligned region CNN (AR-CNN) to tackle the position shift problem. First, a region feature (RF) alignment module with adjacent similarity constraint is designed to consistently predict the position shift between two modalities and adaptively align the cross-modal RFs. Second, we propose a novel region of interest (RoI) jitter strategy to improve the robustness to unexpected shift patterns. Third, we present a new multimodal feature fusion method that selects the more reliable feature and suppresses the less useful one via feature reweighting. In addition, by locating bounding boxes in both modalities and building their relationships, we provide novel multimodal labeling named KAIST-Paired. Extensive experiments on 2-D and 3-D object detection, RGB-T, and RGB-D datasets demonstrate the effectiveness and robustness of our method. Lu Zhang 0054, Zhiyong Liu 0001, Xiangyu Zhu 0001, Zhan Song, Xu Yang 0004, Zhen Lei 0001, Hong Qiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MemoNav: Working Memory Model for Visual NavigationabstractImage-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use all historical observations for decision-making without considering the goal-relevant fraction. To address this limitation, we present MemoNav, a novel memory model for image-goal navigation, which utilizes a working memory-inspired pipeline to improve navigation performance. Specifically, we employ three types of navigation memory. The node features on a map are stored in the short-term memory (STM), as these features are dynamically updated. A forgetting module then retains the informative STM fraction to increase efficiency. We also introduce long-term memory (LTM) to learn global scene representations by progressively aggregating STM features. Subsequently, a graph attention module encodes the retained STM and the LTM to generate working memory (WM) which contains the scene features essential for efficient navigation. The synergy among these three memory types boosts navigation performance by enabling the agent to learn and leverage goal-relevant scene features within a topological map. Our evaluation on multi-goal tasks demonstrates that MemoNav significantly outperforms previous methods across all difficulty levels in both Gibson and Matterport3D scenes. Qualitative results further illustrate that MemoNav plans more efficient routes. Xu Yang 0004, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2024 | VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and ProprioceptionabstractThis paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real entries, collected via simulations in Mujoco and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/. Zhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi 0001, Wang Wei Lee, Minglei Lu, Xiao Teng, Xu Yang 0004, Ming-Hsuan Yang 0001, Hui Cheng 0002 |
ICML | 10 |
| 2024 | Enhancing class-incremental object detection in remote sensing through instance-aware distillation
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
Neurocomputing | 3 |
| 2024 | DeepMulticut: Deep Learning of Multicut Problem for Neuron Segmentation From Electron Microscopy VolumeabstractSuperpixel aggregation is a powerful tool for automated neuron segmentation from electron microscopy (EM) volume. However, existing graph partitioning methods for superpixel aggregation still involve two separate stages-model estimation and model solving, and therefore model error is inherent. To address this issue, we integrate the two stages and propose an end-to-end aggregation framework based on deep learning of the minimum cost multicut problem called DeepMulticut. The core challenge lies in differentiating the NP-hard multicut problem, whose constraint number is exponential in the problem size. With this in mind, we resort to relaxing the combinatorial solver-the greedy additive edge contraction (GAEC)-to a continuous Soft-GAEC algorithm, whose limit is shown to be the vanilla GAEC. Such relaxation thus allows the DeepMulticut to integrate edge cost estimators, Edge-CNNs, into a differentiable multicut optimization system and allows a decision-oriented loss to feed decision quality back to the Edge-CNNs for adaptive discriminative feature learning. Hence, the model estimators, Edge-CNNs, can be trained to improve partitioning decisions directly while beyond the NP-hardness. Also, we explain the rationale behind the DeepMulticut framework from the perspective of bi-level optimization. Extensive experiments on three public EM datasets demonstrate the effectiveness of the proposed DeepMulticut. Zhenchen Li, Xu Yang 0004, Bei Hong, Hao Zhai 0003, Lijun Shen, Xi Chen 0031, Zhiyong Liu 0001, Hua Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Automatically Discovering Novel Visual Categories With Adaptive Prototype LearningabstractThis article targets the task of novel category discovery (NCD), which aims to discover unknown categories when a certain number of classes are already known. The NCD task is challenging due to its closeness to real-world scenarios, where we have only encountered some partial classes and corresponding images. Unlike previous approaches to NCD, we propose a novel adaptive prototype learning method that leverages prototypes to emphasize category discrimination and alleviate the issue of missing annotations for novel classes. Concretely, the proposed method consists of two main stages: prototypical representation learning and prototypical self-training. In the first stage, we develop a robust feature extractor that could effectively handle images from both base and novel categories. This ability of instance and category discrimination of the feature extractor is boosted by self-supervised learning and adaptive prototypes. In the second stage, we utilize the prototypes again to rectify offline pseudo labels and train a final parametric classifier for category clustering. We conduct extensive experiments on four benchmark datasets, demonstrating our method's effectiveness and robustness with state-of-the-art performance. Lu Zhang 0054, Lu Qi 0001, Xu Yang 0004, Hong Qiao, Ming-Hsuan Yang 0001, Zhiyong Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Unseen Object Instance Segmentation with Fully Test-time RGB-D Embeddings AdaptationabstractSegmenting unseen objects is a crucial ability for the robot since it may encounter new environments during the operation. Recently, a popular solution is leveraging RGB-D features of large-scale synthetic data and directly applying the model to unseen real-world scenarios. However, the domain shift caused by the sim2real gap is inevitable, posing a crucial challenge to the segmentation model. In this paper, we em-phasize the adaptation process across sim2real domains and model it as a learning problem on the BatchNorm param-eters of a simulation-trained model. Specifically, we propose a novel non-parametric entropy objective, which formulates the learning objective for the test-time adaptation in an open-world manner. Then, a cross-modality knowledge distillation objective is further designed to encourage the test-time knowledge transfer for feature enhancement. Our approach can be efficiently implemented with only test images, without requiring annotations or revisiting the large-scale synthetic training data. Besides significant time savings, the proposed method consistently improves segmentation results on the overlap and boundary metrics, achieving state-of-the-art performance on unseen object instance segmentation. Lu Zhang 0054, Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
ICRA | 3 |
| 2023 | Incremental Few-Shot Object Detection with scale- and centerness-aware weight generation
Lu Zhang 0054, Xu Yang 0004, Lu Qi 0001, Shaofeng Zeng, Zhiyong Liu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2023 | RTDOD: A large-scale RGB-thermal domain-incremental object detection dataset for UAVs
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
Image Vis. Comput. | 5 |
| 2023 | Static-dynamic global graph representation for pedestrian trajectory prediction
Hao Zhou 0014, Xu Yang 0004, Mingyu Fan, Hai Huang 0004, Dongchun Ren, Huaxia Xia |
Knowl. Based Syst. | 2 |
| 2023 | CSR: Cascade Conditional Variational Auto Encoder with Socially-aware Regression for Pedestrian Trajectory Prediction
Hao Zhou 0014, Dongchun Ren, Xu Yang 0004, Mingyu Fan, Hai Huang 0004 |
Pattern Recognit. | 3 |
| 2023 | CSIR: Cascaded Sliding CVAEs With Iterative Socially-Aware Rethinking for Trajectory PredictionabstractPedestrian trajectory prediction is a hot research topic in many applications, such as video surveillance and autonomous driving. Although many efforts have been done on this topic, there are still many challenges, including accumulated prediction errors, insufficient training data usage, and future-past incompatibility. To overcome these challenges, we propose a novel trajectory prediction method, called CSIR, which consists of a cascaded sliding conditional variational autoencoder (CS-CVAE) module and an iterative future-past social compatible rethinking (I-SCR) module. The CS-CVAE module reduces the accumulated prediction errors by using cascaded prediction models for the early future time steps. In this way, the training losses of the early time steps are separately considered and minimized from the later losses. For the following time steps in CS-CVAE, a sliding prediction model with a longer observation time span is used and additional data from the future time span can be collected for training. On the other hand, the I-SCR module generates offsets to improve the predictions iteratively by checking the interaction compatibility between the predicted trajectories and the past trajectories, which resembles with the human rethinking mechanism in motion planning. Experiments results on two widely explored pedestrian trajectory prediction datasets, Stanford Drone Dataset (SDD) and ETH/UCY, show that the proposed method surpasses previous state-of-the-art methods by notable margins. Hao Zhou 0014, Xu Yang 0004, Dongchun Ren, Hai Huang 0004, Mingyu Fan |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | A Learning Method for Feature Correspondence with OutliersabstractFeature correspondence is an important topic in many computer vision or robot vision tasks. Different from traditional optimization based matching method, in the last two years, researchers are finally able to solve the matching process in a learning manner. As a representative method, SuperGlue achieves superior performance in many real-world tasks, but it still has problems in dealing with outlier features. Targeting at the outlier problem, this paper improves SuperGlue by introducing a deep learning based feature correspondence method, which consists of the pruned attentional graph neural network and the improved matching layer for the outlier problem. Experiments on real world images validate the effectiveness of the proposed method. Xu Yang 0004, Shaofeng Zeng, Yu-Chen Lu, Zhiyong Liu 0001 |
ICPR | 1 |
| 2022 | Graph matching based on fast normalized cut and multiplicative update mapping
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001 |
Pattern Recognit. | 2 |
| 2022 | CANet: Co-attention network for RGB-D semantic segmentation
Hao Zhou 0014, Lu Qi 0001, Hai Huang 0004, Xu Yang 0004, Zhaoliang Wan, Xianglong Wen |
Pattern Recognit. | 4 |
| 2022 | Incremental few-shot object detection via knowledge transfer
Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
Pattern Recognit. Lett. | 3 |
| 2021 | Graph matching based point correspondence with alternating direction method of multipliers
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001, Mingyu Fan |
Neurocomputing | 2 |
| 2021 | AST-GNN: An attention-based spatio-temporal graph neural network for Interaction-aware pedestrian trajectory prediction
Hao Zhou 0014, Dongchun Ren, Huaxia Xia, Mingyu Fan, Xu Yang 0004, Hai Huang 0004 |
Neurocomputing | 5 |
| 2021 | Supervised learning for parameterized Koopmans-Beckmann's graph matching
Shaofeng Zeng, Zhiyong Liu 0001, Xu Yang 0004 |
Pattern Recognit. Lett. | 3 |
| 2021 | A Doubly Graduated Method for Inference in Markov Random FieldabstractMaximum a posteriori (MAP) inference in Markov random field (MRF) lays the foundation for many computer vision tasks, which can be formulated by a binary quadratic programming (BQP) problem. Compared with the discrete methods, the continuous relaxation scheme becomes popular due to its generality and efficiency. However, existing continuous relaxation based MAP algorithms are still limited by two problems, i.e., the highly nonconvex original objective function and the gap between the original BQP problem and the relaxed continuous optimization problem. Targeting the two problems, this paper presents a doubly graduated continuous relaxation algorithm for MAP inference in MRF, which are, respectively, the Gaussian smoothing based graduated nonconvexity process and conditional gradient ascent based graduated projection. Experiments on both synthetic data and real-world images illustrate the algorithm's state-of-the-art performance in objective function optimization and typical computer vision tasks. Xu Yang 0004, Zhiyong Liu 0001 |
SIAM J. Imaging Sci. | 1 |
| 2020 | RGB-D Co-attention Network for Semantic Segmentation
Hao Zhou 0014, Lu Qi 0001, Zhaoliang Wan, Hai Huang 0004, Xu Yang 0004 |
ACCV (1) | 5 |
| 2020 | GSDCN: A Customized Two-Stage Neural Network for Benthonic Organism Detection
Zhaoliang Wan, Lu Zhang 0054, Hai Huang 0004, Xu Yang 0004 |
ICONIP (2) | 4 |
| 2020 | Knowledge-Experience Graph with Denoising Autoencoder for Zero-Shot Learning in Visual Cognitive Development
Xu Yang 0004, Zhiyong Liu 0001, Lu Zhang 0054, Dongchun Ren, Mingyu Fan |
ICONIP (5) | 2 |
| 2020 | An Attention-Based Interaction-Aware Spatio-Temporal Graph Neural Network for Trajectory Prediction
Hao Zhou 0014, Dongchun Ren, Huaxia Xia, Mingyu Fan, Xu Yang 0004, Hai Huang 0004 |
ICONIP (5) | 5 |
| 2020 | MixedFusion: 6D Object Pose Estimation from Decoupled RGB-Depth FeaturesabstractEstimating the 6D pose of objects is an important process for intelligent systems to achieve interaction with the real-world. As the RGB-D sensors become more accessible, the fusion-based methods have prevailed, since the point clouds provide complementary geometric information with RGB values. However, due to the difference in feature space between color image and depth image, the network structures that directly perform point-to-point matching fusion do not effectively fuse the features of the two. In this paper, we propose a simple but effective approach, named MixedFusion. Different from the prior works, we argue that the spatial correspondence of color and point clouds could be decoupled and reconnected, thus enabling a more flexible fusion scheme. By performing the proposed method, more informative points can be mixed and fused with rich color features. Extensive experiments are conducted on the challenging LineMod and YCB-Video datasets, which shows that our method significantly boosts the performance without introducing extra overheads. Furthermore, when the minimum tolerance of metric narrows, the proposed approach performs better for the high-precision demands. Hangtao Feng, Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001 |
ICPR | 3 |
| 2020 | Activity and Relationship Modeling Driven Weakly Supervised Object DetectionabstractThis paper presents a weakly supervised object detection method based on activity label and relationship modeling, which is motivated by the assumption that configuration of human and object are similar in same activity, and joint modeling of human, active object and activity could leverage the recognition of them. Compared to most weakly supervised method taking object as independent instance, firstly, active human and object proposals are learned and filtered based on class activation map of multi-label classification. Secondly, a spatial relationship prior including relative position, scale, overlaps etc are learned dependent on action. Finally, a multi-stream object detection framework integrating the spatial prior and pairwise ROI pooling are proposed to jointly learn the object and action class. Experiments are conducted on HICO-DET dataset, and our approach outperforms the state of the art weakly supervised object detection methods. Yinlin Li, Xu Yang 0004, Yuren Zhang |
ICPR | 3 |
| 2020 | Encoding primitives generation policy learning for robotic arm to overcome catastrophic forgetting in sequential multi-tasks learning
Fangzhou Xiong, Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Hong Qiao, Amir Hussain 0001 |
Neural Networks | 4 |
| 2020 | A Continuation Method for Graph Matching Based Feature CorrespondenceabstractFeature correspondence lays the foundation for many computer vision and image processing tasks, which can be well formulated and solved by graph matching. Because of the high complexity, approximate methods are necessary for graph matching, and the continuous relaxation provides an efficient approximate scheme. But there are still many problems to be settled, such as the highly nonconvex objective function, the ignorance of the combinatorial nature of graph matching in the optimization process, and few attention to the outlier problem. Focusing on these problems, this paper introduces a continuation method directly targeting at the combinatorial optimization problem associated with graph matching. Specifically, first a regularization function incorporating the original objective function and the discrete constraints is proposed. Then a continuation method based on Gaussian smoothing is applied to it, in which the closed forms of relevant functions with respect to the outlier distribution are deduced. Experiments on both synthetic data and real world images validate the effectiveness of the proposed method. Xu Yang 0004, Zhiyong Liu 0001, Hong Qiao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Incorporating Discrete Constraints Into Random Walk-Based Graph MatchingabstractGraph matching is a fundamental problem in theoretical computer science and artificial intelligence, and lays the foundation for many computer vision and machine learning tasks. Approximate algorithms are necessary for graph matching due to its NP-complete nature. Inspired by the usage in network-related tasks, random walk is generalized to graph matching as a type of approximate algorithm. However, it may be inappropriate for the previous random walk-based graph matching algorithms to utilize continuous techniques without considering the discrete property. In this paper, we propose a novel random walk-based graph matching algorithm by incorporating both continuous and discrete constraints in the optimization process. Specifically, after interpreting graph matching by random walk, the continuous constraints are directly embedded in the random walk constraint in each iteration. Further, both the assignment matrix (vector) and the pairwise similarity measure between graphs are iteratively updated according the discrete constraints, which automatically leads the continuous solution to the discrete domain. Comparisons on both synthetic and real-world data demonstrate the effectiveness of the proposed algorithm. Xu Yang 0004, Zhiyong Liu 0001, Hong Qiaoxu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not strictly aligned, making one object has different positions in different modalities. In deep learning based methods, this problem makes it difficult to fuse the feature maps from both modalities and puzzles the CNN training. In this paper, we propose a novel Aligned Region CNN (AR-CNN) to handle the weakly aligned multispectral data in an end-to-end way. Firstly, we design a Region Feature Alignment (RFA) module to capture the position shift and adaptively align the region features of the two modalities. Secondly, we present a new multimodal fusion method, which performs feature re-weighting to select more reliable features and suppress the useless ones. Besides, we propose a novel RoI jitter strategy to improve the robustness to unexpected shift patterns of different devices and system settings. Finally, since our method depends on a new kind of labelling: bounding boxes that match each modality, we manually relabel the KAIST dataset by locating bounding boxes in both modalities and building their relationships, providing a new KAIST-Paired Annotation. Extensive experimental validations on existing datasets are performed, demonstrating the effectiveness and robustness of the proposed method. Code and data are available at https://github.com/luzhang16/AR-CNN. Lu Zhang 0054, Xiangyu Zhu 0001, Xu Yang 0004, Zhen Lei 0001, Zhiyong Liu 0001 |
ICCV | 4 |
| 2019 | Learning Transferable Policies with Improved Graph Neural Networks on Serial Robotic Structure
Fangzhou Xiong, Xu Yang 0004, Zhiyong Liu 0001 |
ICONIP (3) | 3 |
| 2019 | Sub-hypergraph matching based on adjacency tensor
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | Faster R-CNN for marine organisms detection and recognition using data augmentation
Hai Huang 0004, Hao Zhou 0014, Xu Yang 0004, Lu Zhang 0054, Lu Qi 0001, Ai-Yun Zang |
Neurocomputing | 3 |
| 2019 | Special issue on advances in graph algorithm and applications
Zhiyong Liu 0001, Kaizhu Huang, Xu Yang 0004, Cheng-Lin Liu 0001 |
Neurocomputing | 3 |
| 2019 | Guided Policy Search for Sequential Multitask LearningabstractPolicy search in reinforcement learning (RL) is a practical approach to interact directly with environments in parameter spaces, that often deal with dilemmas of local optima and real-time sample collection. A promising algorithm, known as guided policy search (GPS), is capable of handling the challenge of training samples using trajectory-centric methods. It can also provide asymptotic local convergence guarantees. However, in its current form, the GPS algorithm cannot operate in sequential multitask learning scenarios. This is due to its batch-style training requirement, where all training samples are collectively provided at the start of the learning process. The algorithm’s adaptation is thus hindered for real-time applications, where training samples or tasks can arrive randomly. In this paper, the GPS approach is reformulated, by adapting a recently proposed, lifelong-learning method, and elastic weight consolidation. Specifically, Fisher information is incorporated to impart knowledge from previously learned tasks. The proposed algorithm, termed sequential multitask learning-GPS, is able to operate in sequential multitask learning settings and ensuring continuous policy learning, without catastrophic forgetting. Pendulum and robotic manipulation experiments demonstrate the new algorithms efficacy to learn control policies for handling sequentially arriving training samples, delivering comparable performance to the traditional, and batch-based GPS algorithm. In conclusion, the proposed algorithm is posited as a new benchmark for the real-time RL and robotics research community. Fangzhou Xiong, Biao Sun 0005, Xu Yang 0004, Hong Qiao, Kaizhu Huang, Amir Hussain 0001, Zhiyong Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | Overcoming Catastrophic Forgetting with Self-adaptive Identifiers
Fangzhou Xiong, Zhiyong Liu 0001, Xu Yang 0004 |
ICONIP (3) | 3 |
| 2018 | Graph Matching Based on Fast Normalized Cut
Jing Yang 0067, Xu Yang 0004, Zhangbing Zhou, Zhiyong Liu 0001 |
ICONIP (6) | 2 |
| 2018 | Single Shot Feature Aggregation Network for Underwater Object DetectionabstractThe rapidly developing ocean exploration and observation make the demand for underwater object detection become increasingly urgent. Recently, deep convolutional neural networks (CNN) have shown strong ability in feature representation and CNN-based detectors also achieve remarkable performance, but still facing the big challenge when detecting multi-scale objects in a complex underwater environment. To address this challenge, we propose a novel underwater object detector, introducing multiscale features and complementary context information for better classification and location ability. In the auto-grabbing contest of 2017 Underwater Robot Picking Contest sponsored by National Natural Science Foundation of China (NSFC), we won the 1-st place by using proposed method for real coastal underwater object detection. Lu Zhang 0054, Xu Yang 0004, Zhiyong Liu 0001, Lu Qi 0001, Hao Zhou 0014, Charles Chiu |
ICPR | 2 |
| 2018 | Adaptive Graph MatchingabstractEstablishing correspondence between point sets lays the foundation for many computer vision and pattern recognition tasks. It can be well defined and solved by graph matching. However, outliers may significantly deteriorate its performance, especially when outliers exist in both point sets and meanwhile the inlier number is unknown. In this paper, we propose an adaptive graph matching algorithm to tackle this problem. Specifically, a novel formulation is proposed to make the graph matching model adaptively determine the number of inliers and match them, then by relaxing the discrete domain to its convex hull the discrete optimization problem is relaxed to be a continuous one, and finally a graduated projection scheme is used to get a discrete matching solution. Consequently, the proposed algorithm could realize inlier number estimation, inlier selection, and inlier matching in one optimization framework. Experiments on both synthetic data and real world images witness the effectiveness of the proposed algorithm. Xu Yang 0004, Zhiyong Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | An Algorithm for Finding the Most Similar Given Sized Subgraphs in Two Weighted GraphsabstractWe propose a weighted common subgraph (WCS) matching algorithm to find the most similar subgraphs in two labeled weighted graphs. WCS matching, as a natural generalization of equal-sized graph matching and subgraph matching, has found wide applications in many computer vision and machine learning tasks. In this brief, WCS matching is first formulated as a combinatorial optimization problem over the set of partial permutation matrices. Then, it is approximately solved by a recently proposed combinatorial optimization framework-graduated nonconvexity and concavity procedure. Experimental comparisons on both synthetic graphs and real-world images validate its robustness against noise level, problem size, outlier number, and edge density. Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | A Linear Online Guided Policy Search Algorithm
Biao Sun 0005, Fangzhou Xiong, Zhiyong Liu 0001, Xu Yang 0004, Hong Qiao |
ICONIP (5) | 4 |
| 2017 | A Bayesian Posterior Updating Algorithm in Reinforcement Learning
Fangzhou Xiong, Zhiyong Liu 0001, Xu Yang 0004, Biao Sun 0005, Charles Chiu, Hong Qiao |
ICONIP (5) | 3 |
| 2017 | LibCoopt: A library for combinatorial optimization on partial permutation matrices
Zhiyong Liu 0001, Xu Yang 0004, Shi-Hao Feng |
Neurocomputing | 3 |
| 2017 | Probabilistic hypergraph matching based on affinity tensor updating
Xu Yang 0004, Zhiyong Liu 0001, Hong Qiao, Jianhua Su |
Neurocomputing | 1 |
| 2017 | Point correspondence by a new third order graph matching algorithm
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
Pattern Recognit. | 1 |
| 2016 | Stitching contaminated images
Chuan Li 0004, Zhiyong Liu 0001, Xu Yang 0004, Hong Qiao, Jianhua Su |
Neurocomputing | 3 |
| 2016 | Introducing locally affine-invariance constraints into lunar surface image correspondence
Yuren Zhang, Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001, Chuankai Liu |
Neurocomputing | 2 |
| 2015 | Feature correspondence based on directed structural model matching
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
Image Vis. Comput. | 1 |
| 2015 | Outlier robust point correspondence based on GNCCP
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | Graph Matching by Simplified Convex-Concave Relaxation Procedure
Zhiyong Liu 0001, Hong Qiao, Xu Yang 0004, Steven C. H. Hoi |
Int. J. Comput. Vis. | 3 |
| 2013 | Partial correspondence based on subgraph matching
Xu Yang 0004, Hong Qiao, Zhiyong Liu 0001 |
Neurocomputing | 1 |