VLDB 2026 Research / reviewers in the wild / expert
Sixian Zhang
dblp:251/1108
· DBLP profile ↗
20ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-1065-5348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning on the Go: A Meta-Learning Object Navigation Model
Xiaorong Qin, Xinhang Song, Sixian Zhang, Xinyao Yu 0002, Xinmiao Zhang 0004, Shuqiang Jiang |
ICCV | 3 |
| 2025 | Function-Centric Bayesian Network for Zero-Shot Object Goal Navigation
Sixian Zhang, Xinyao Yu 0002, Xinhang Song, Yiyao Wang, Shuqiang Jiang |
ICCV | 1 |
| 2025 | HSI and LiDAR Joint Classification Network With Interactive Extraction and Heterogeneous BalanceabstractTo improve the representation of spectral features and enhance the complementarity of heterogeneous spatial features in the joint classification of hyperspectral image (HSI) and light detection and ranging image (LiDAR), this paper proposes a joint classification network with interactive extraction and heterogeneous balance (JCNIH). On the one hand, this paper designs an interactive method for information exchange between large-span spectral features and small-span spectral features, which improves the adaptability and representation of spectral features. On the other hand, this paper designs a balance loss to constrain the extraction of heterogenous spatial features, which can keep the feature extraction balance and enhance the spatial complementarity. Then, the extracted spectral features and heterogeneous spatial features are aggregated for classification. Finally, comparative and ablation experiments demonstrate the effectiveness of JCNIH, and the experimental results show that the JCNIH outperforms the state-of-the-art methods in overall accuracy by 10.33% on three public datasets. Pengbo Mi, Yi Yang 0008, Meng Zhang 0029, Sixian Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | HOZ++: Versatile Hierarchical Object-to-Zone Graph for Object NavigationabstractThe goal of object navigation task is to reach the expected objects using visual information in unseen environments. Previous works typically implement deep models as agents that are trained to predict actions based on visual observations. Despite extensive training, agents often fail to make wise decisions when navigating in unseen environments toward invisible targets. In contrast, humans demonstrate a remarkable talent to navigate toward targets even in unseen environments. This superior capability is attributed to the cognitive map in the hippocampus, which enables humans to recall past experiences in similar situations and anticipate future occurrences during navigation. It is also dynamically updated with new observations from unseen environments. The cognitive map equips humans with a wealth of prior knowledge, significantly enhancing their navigation capabilities. Inspired by human navigation mechanisms, we propose the Hierarchical Object-to-Zone (HOZ++) graph, which encapsulates the regularities among objects, zones, and scenes. The HOZ++ graph helps the agent to identify the current zone and the target zone, and computes an optimal path between them, then selects the next zone along the path as the guidance for the agent. Moreover, the HOZ++ graph continuously updates based on real-time observations in new environments, thereby enhancing its adaptability to new environments. Our HOZ++ graph is versatile and can be integrated into existing methods, including end-to-end RL and modular methods. Our method is evaluated across four simulators, including AI2-THOR, RoboTHOR, Gibson, and Matterport 3D. Additionally, we build a realistic environment to evaluate our method in the real world. Experimental results demonstrate the effectiveness and efficiency of our proposed method. Sixian Zhang, Xinhang Song, Xinyao Yu 0002, Yubing Bai, Xinlong Guo, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | An Interactive Navigation Method with Effect-oriented AffordanceabstractVisual navigation is to let the agent reach the target according to the continuous visual input. In most previous works, visual navigation is usually assumed to be done in a static and ideal environment: the target is always reachable with no need to alter the environment. However, the “messy” environments are more general and practical in our daily lives, where the agent may get blocked by obstacles. Thus Interactive Navigation (InterNav) is introduced to navigate to the objects in more realistic “messy” environments according to the object interaction. Prior work on InterNav learns shortterm interaction through extensive trials with reinforcement learning. However, interaction does not guarantee efficient navigation, that is, plan-ning obstacle interactions that make shorter paths and con-sume less effort is also crucial. In this paper, we introduce an effect-oriented affordance map to enable longterm interactive navigation, extending the existing map-based nav-igation framework to the domain of dynamic environment. We train a set of affordance functions predicting available interactions and the time cost of removing obstacles, which informatively support an interactive modular system to ad-dress interaction and longterm planning. Experiments on the ProcTHOR simulator demonstrate the capability of our affordance-driven system in longterm navigation in complex dynamic environments. Yuehu Liu, Xinhang Song, Yuyi Liu, Sixian Zhang, Shuqiang Jiang |
CVPR | 5 |
| 2024 | Learning Multi-Dimensional Human Preference for Text-to-Image GenerationabstractCurrent metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images, they reduce the rich tapestry of human preference to a single overall score. However, the preference results vary when humans evaluate images with different aspects. Therefore, to learn the multidimensional human preferences, we propose the Multi-dimensional Preference Score (MPS), the first multidimensional preference scoring model for the evaluation of text-to-image models. The MPS introduces the preference condition module upon CLIP model to learn these diverse preferences. It is trained based on our Multi-dimensional Human Preference (MHP) Dataset, which comprises 918,315 human preference choices across four dimensions (i.e., aesthetics, semantic alignment, detail quality and overall assessment) on 607,541 images. The images are generated by a wide range of latest text-to-image models. The MPS outperforms existing scoring methods across 3 datasets in 4 dimensions, enabling it a promising metric for evaluating and improving text-to-image generation. The model and dataset will be made publicly available to facilitate future research. Project page: htt ps: //wangbohan97.github.io/MPS/. Sixian Zhang, Junqiang Wu, Yan Li 0043, Tingting Gao, Di Zhang 0026, Zhongyuan Wang 0006 |
CVPR | 1 |
| 2024 | Imagine Before Go: Self-Supervised Generative Map for Object Goal NavigationabstractThe Object Goal navigation (ObjectNav) task requires the agent to navigate to a specified target in an unseen environment. Since the environment layout is unknown, the agent needs to infer the unknown contextual objects from partially observations, thereby deducing the likely location of the target. Previous end-to-end RL methods capture contextual relationships through implicit representations while they lack notion of geometry. Alternatively, modular methods construct local maps for recording the observed geometric structure of unseen environment, however, lacking the reasoning of contextual relation limits the exploration efficiency. In this work, we propose the self-supervised generative map (SGM), a modular method that learns the explicit context relation via self-supervised learning. The SGM is trained to leverage both episodic observations and general knowledge to reconstruct the masked pixels of a cropped global map. During navigation, the agent maintains an incomplete local semantic map, meanwhile, the unknown regions of the local map are generated by the pretrained SGM. Based on the generated map, the agent sets the predicted location of the target as the goal and moves towards it. Experiments on Gibson, MP3D and HM3D show the effectiveness of our method. The code is available at https://github.com/sx-zhang/SGM. Sixian Zhang, Xinyao Yu 0002, Xinhang Song, Shuqiang Jiang |
CVPR | 1 |
| 2024 | Foreground Aware Correlation Filter with Adaptive Feature Response Fusion for Real-Time UAV TrackingabstractBackground Aware Correlation Filter (BACF) tracker achieves accurate tracking result in visual object tracking by mitigating boundary effects, yet is limited in challenging scenarios especially in viewpoint change and illumination variation, which are frequently encountered in Unmanned Aerial Vehicle (UAV) tracking tasks. To address the shortcomings, we propose a Foreground Aware Correlation Filter with adaptive feature response fusion (FACF). In this paper, we use saliency detection to generate foreground prior knowledge in training phase for suppressing potential noise. Furthermore, recognizing the limitation of BACF, which relies on a single feature, a novel adaptive fusion strategy is designed to fuse multiple feature responses during the detection phase. This strategy aims to enhance the robustness of the tracker. Extensive experiments have been conducted on three challenging benchmarks. The tracking results show that the proposed tracker performs accurate and robust tracking result and satisfies real-time requirement with 48.28fps. Zhuo Xiao, Yi Yang 0008, Sixian Zhang, Wenbiao Li, Pengrong Bao, Deqiang Han |
FUSION | 3 |
| 2024 | Trajectory Diffusion for ObjectGoal NavigationabstractObject goal navigation requires an agent to navigate to a specified object in an unseen environment based on visual observations and user-specified goals.
Human decision-making in navigation is sequential, planning a most likely sequence of actions toward the goal.
However, existing ObjectNav methods, both end-to-end learning methods and modular methods, rely on single-step planning. They output the next action based on the current model input, which easily overlooks temporal consistency and leads to myopic planning.
To this end, we aim to learn sequence planning for ObjectNav. Specifically, we propose trajectory diffusion to learn the distribution of trajectory sequences conditioned on the current observation and the goal.
We utilize DDPM and automatically collected optimal trajectory segments to train the trajectory diffusion.
Once the trajectory diffusion model is trained, it can generate a temporally coherent sequence of future trajectory for agent based on its current observations.
Experimental results on the Gibson and MP3D datasets demonstrate that the generated trajectories effectively guide the agent, resulting in more accurate and efficient navigation. Xinyao Yu 0002, Sixian Zhang, Xinhang Song, Xiaorong Qin, Shuqiang Jiang |
NeurIPS | 2 |
| 2023 | Layout-based Causal Inference for Object NavigationabstractPrevious works for ObjectNav task attempt to learn the association (e.g. relation graph) between the visual inputs and the goal during training. Such association contains the prior knowledge of navigating in training environments, which is denoted as the experience. The experience performs a positive effect on helping the agent infer the likely location of the goal when the layout gap between the unseen environments of the test and the prior knowledge obtained in training is minor. However, when the layout gap is significant, the experience exerts a negative effect on navigation. Motivated by keeping the positive effect and removing the negative effect of the experience, we propose the layout-based soft Total Direct Effect (L-sTDE) framework based on the causal inference to adjust the prediction of the navigation policy. In particular, we propose to calculate the layout gap which is defined as the KL divergence between the posterior and the prior distribution of the object layout. Then the sTDE is proposed to appropriately control the effect of the experience based on the layout gap. Experimental results on AI2THOR, RoboTHOR, and Habitat demonstrate the effectiveness of our method. The code is available at https://github.com/sx-zhang/Layout-based-sTDE.git. Sixian Zhang, Xinhang Song, Yubing Bai, Xinyao Yu 0002, Shuqiang Jiang |
CVPR | 1 |
| 2023 | A Variational Method with Kernel Estimation and Low Rank Prior for PansharpeningabstractIn this article, a new variational pansharpening method based on kernel estimation and regional extended low rank is proposed, which aims to generate a high resolution multispectral (HRMS) image by fusing the panchromatic (PAN) and multispectral (MS) image. First, an estimated blurring kernel is generated for the spectral constraint term, which can build the relationship between the MS and HRMS image more accurately and improve the spectral quality of the HRMS image. Second, a spatial constrain term is designed by adopting the proportional relationship of the PAN and HRMS image in gradient domain, which preserves the geometric information of the PAN image well. Third, according to sensor imaging principle, a prior constraint term is proposed based on regional extended low rank, which can improve the spatial clarity of HRMS image. The above three constraint terms are combined to form the proposed variational pansharpening method, and the ADMM method is applied for solving it. Finally, experiments show the effectiveness of the proposed method through comparing with other state-of-art pansharpening methods. Pengbo Mi, Yi Yang 0008, Meng Zhang 0029, Sixian Zhang, Erqi Zhang, Wenbiao Li |
FUSION | 4 |
| 2023 | Long-Short Term Policy for Visual Object NavigationabstractThe goal of visual object navigation for an agent is to find the target objects accurately. Recent works mainly focus on the feature of embedding, attempting to learn better features with different variants, such as object distribution and graph representations. However, some typical navigation problems in complex environments, such as partially known and obstacle problems, may not be effectively addressed by previous feature embedding methods. In this paper, we propose a framework with a long-short objective policy, where the hidden states are classified according to the navigation objectives at that moment and separately rewarded. Specifically, we consider two objectives: the long-term objective is to go closer to the target, and the short-term objective is for obstacle avoidance and exploration. To alleviate the effect of long-term and short-term alternation, we build a state memory and propose an adjustment gate to update the state memory. Finally, all past hidden states are reweighted and combined for action prediction with an action-boosting gate. Experimental results on RoboTHOR show that the proposed method can significantly outperform the state-of-the-art. Yubing Bai, Xinhang Song, Sixian Zhang, Shuqiang Jiang |
IROS | 4 |
| 2022 | Generative Meta-Adversarial Network for Unseen Object Navigation
Sixian Zhang, Xinhang Song, Yubing Bai, Shuqiang Jiang |
ECCV (39) | 1 |
| 2021 | A Mutli-feature Correlation Filter Tracker with Different Hash Algorithm
Sixian Zhang, Yi Yang 0008, Meng Zhang 0029, Pengbo Mi |
FUSION | 1 |
| 2021 | Hierarchical Object-to-Zone Graph for Object NavigationabstractThe goal of object navigation is to reach the expected objects according to visual information in the unseen environments. Previous works usually implement deep models to train an agent to predict actions in real-time. However, in the unseen environment, when the target object is not in egocentric view, the agent may not be able to make wise decisions due to the lack of guidance. In this paper, we propose a hierarchical object-to-zone (HOZ) graph to guide the agent in a coarse-to-fine manner, and an online-learning mechanism is also proposed to update HOZ according to the real-time observation in new environments. In particular, the HOZ graph is composed of scene nodes, zone nodes and object nodes. With the pre-learned HOZ graph, the real-time observation and the target goal, the agent can constantly plan an optimal path from zone to zone. In the estimated path, the next potential zone is regarded as sub-goal, which is also fed into the deep reinforcement learning model for action prediction. Our methods are evaluated on the AI2-Thor simulator. In addition to widely used evaluation metrics SR and SPL, we also propose a new evaluation metric of SAE that focuses on the effective action rate. Experimental results demonstrate the effectiveness and efficiency of our proposed method. The code is available at https://github.com/sx-zhang/HOZ.git. Sixian Zhang, Xinhang Song, Yubing Bai, Yakui Chu, Shuqiang Jiang |
ICCV | 1 |
| 2021 | ION: Instance-level Object NavigationabstractVisual object navigation is a fundamental task in Embodied AI. Previous works focus on the category-wise navigation, in which navigating to any possible instance of target object category is considered a success. Those methods may be effective to find the general objects. However, it may be more practical to navigate to the specific instance in our real life, since our particular requirements are usually satisfied with specific instances rather than all instances of one category. How to navigate to the specific instance has been rarely researched before and is typically challenging to current works. In this paper, we introduce a new task of Instance Object Navigation (ION), where instance-level descriptions of targets are provided and instance-level navigation is required. In particular, multiple types of attributes such as colors, materials and object references are involved in the instance-level descriptions of the targets. In order to allow the agent to maintain the ability of instance navigation, we propose a cascade framework with Instance-Relation Graph (IRG) based navigator and instance grounding module. To specify the different instances of the same object categories, we construct instance-level graph instead of category-level one, where instances are regarded as nodes, encoded with the representation of colors, materials and locations (bounding boxes). During navigation, the detected instances can activate corresponding nodes in IRG, which are updated with graph convolutional neural network (GCNN). The final instance prediction is obtained with the grounding module by selecting the candidates (instances) with maximum probability (a joint probability of category, color and material, obtained by corresponding regressors with softmax). For the task evaluation, we build a benchmark for instance-level object navigation on AI2-Thor simulator, where over 27,735 object instance descriptions and navigation groundtruth are automatically obtained through the interaction with the simulator. The proposed model outperforms the baseline in instance-level metrics, showing that our proposed graph model can guide instance object navigation, as well as leaving promising room for further improvement. The project is available at https://github.com/LWJ312/ION. Xinhang Song, Yubing Bai, Sixian Zhang, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2020 | Generalized Zero-shot Learning with Multi-source Semantic Embeddings for Scene RecognitionabstractRecognizing visual categories from semantic descriptions is a promising way to extend the capability of a visual classifier beyond the concepts represented in the training data (i.e. seen categories). This problem is addressed by (generalized) zero-shot learning methods (GZSL), which leverage semantic descriptions that connect them to seen categories (e.g. label embedding, attributes). Conventional GZSL are designed mostly for object recognition. In this paper we focus on zero-shot scene recognition, a more challenging setting with hundreds of categories where their differences can be subtle and often localized in certain objects or regions. Conventional GZSL representations are not rich enough to capture these local discriminative differences. Addressing these limitations, we propose a feature generation framework with two novel components: 1) multiple sources of semantic information (i.e. attributes, word embeddings and descriptions), 2) region descriptions that can enhance scene discrimination. To generate synthetic visual features we propose a two-step generative approach, where local descriptions are sampled and used as conditions to generate visual features. The generated features are then aggregated and used together with real features to train a joint classifier. In order to evaluate the proposed method, we introduce a new dataset for zero-shot scene recognition with multi-semantic annotations. Experimental results on the proposed dataset and SUN Attribute dataset illustrate the effectiveness of the proposed method. Xinhang Song, Haitao Zeng, Sixian Zhang, Luis Herranz, Shuqiang Jiang |
ACM Multimedia | 3 |
| 2019 | Simultaneous Multiple Features Tracking of Beats: A Representation Learning Approach to Reduce False Alarm Rate in ICUsabstractThe high rate of false alarms is a key challenge related to patient care in intensive care units (ICUs) that can result in delayed responses of the medical staff. Several rule-based and machine learning-based techniques have been developed to address this problem. However, the majority of these methods rely on the availability of different physiological signals such as different electrocardiogram (ECG) leads, arterial blood pressure (ABP), and photoplethysmogram (PPG), where each signal is analyzed by an independent processing unit and the results are fed to an algorithm to determine an alarm. That calls for novel methods that can accurately detect the cardiac events by only accessing one signal (e.g., ECG) with a low level of computation and sensors requirement. We propose a novel and robust representation learning framework for ECG analysis that only rely on a single lead ECG signal and yet achieves considerably better performance compared to the state-of-the-art works in this domain, without relying on an expert knowledge. We evaluate the performance of this method using the "2015 Physionet computing in cardiology challenge" dataset. To the best of our knowledge, the best previously reported performance is based on both expert knowledge and machine learning where all available signals of ECG, ABP and PPG are utilized. Our proposed method reaches the performance of 97.3%, 95.5 %, and 90.8 % in terms of sensitivity, specificity, and the challenge's score, respectively for the detection of five arrhythmias when only one single ECG lead signals is used without any expert knowledge. Behzad Ghazanfari, Sixian Zhang, Fatemeh Afghah, Nathan Payton-McCauslin |
BIBM | 2 |
| 2019 | A Real-Time Scene Recognition System Based on RGB-D Video StreamsabstractDepth data captured by the cameras such as Microsoft Kinect can bring depth information than traditional RGB data, which is also more robust to different environments, such as dim or dark lighting conditions. In this technical demonstration, we build a scene recognition system based on real-time processing of RGB-D video streams. Our system recognizes the scenes with video clips, where three types of threads are implemented to ensure the realtime. This system first buffers the frames of both RGB and depth videos with the capturing threads. When the buffered videos reach the certain length, the frames will be packed into clips and forwarded in a pre-trained C3D model to predict scene labels with the scene recognition thread. Finally, the predicted scene labels and captured videos are illustrated in our user interface with illustration thread. Yuyun Hua, Sixian Zhang, Xinhang Song, Jia'ning Li, Shuqiang Jiang |
ICMI | 2 |
| 2019 | Aberrance-aware Gradient-sensitive Attentions for Scene Recognition with RGB-D VideosabstractWith the developments of deep learning, previous approaches have made successes in scene recognition with massive RGB data obtained from the ideal environments. However, scene recognition in real world may face various types of aberrant conditions caused by different unavoidable factors, such as the lighting variance of the environments and the limitations of cameras, which may damage the performance of previous models. In addition to ideal conditions, our motivation is to investigate researches on robust scene recognition models for unconstrained environments. In this paper, we propose an aberrance-aware framework for RGB-D scene recognition, where several types of attentions, such as temporal, spatial and modal attentions are integrated to spatio-temporal RGB-D CNN models to avoid the interference of RGB frame blurring, depth missing, and light variance. All the attentions are homogeneously obtained by projecting the gradient-sensitive maps of visual data into corresponding spaces. Particularly, the gradient maps are captured with the convolutional operations with the typically designed kernels, which can be seamlessly integrated into end-to-end CNN training. The experiments under different challenging conditions demonstrate the effectiveness of the proposed method. Xinhang Song, Sixian Zhang, Yuyun Hua, Shuqiang Jiang |
ACM Multimedia | 2 |