EDBT 2026 Demo / reviewers in the wild / expert
Qiyu Sun
dblp:46/3606
· DBLP profile ↗
24ranked-venue papers
2as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Boundary-Based Active Domain Adaptation for Semantic Segmentation Under Adverse ConditionsabstractExisting domain adaptation semantic segmentation (DASS) methods under adverse conditions often depend on pseudo-labels for network training. However, these pseudo-labels are frequently plagued by noise and bias toward high-confidence predictions, thereby impeding the enhancement of segmentation performance. This article tackles the above challenge by proposing a novel boundary-based active domain adaptation (ADA) framework, which efficiently selects both informative low-confidence samples and high-confident but misclassified samples to be labeled while maximizing the segmentation performance under a limited annotation budget. For the evaluation of sample confidence and informativeness, we first propose ranking weighted feature space impurity (RWFSI) metric to quantify category distribution among a sample's nearest neighbors within the feature space and consider the samples with higher RWFSI values as low-confidence samples around the decision boundary, which can also alleviate the category imbalance of active labels. Subsequently, we apply Gaussian mixture models (GMMs) to model the distribution across source and target domains. Using the spatial arrangement of each GMM component, we define the intraclass domain shift score (ICDSS), which identifies samples with high ICDSS values as those more likely to be high-confidence but misclassified, aiding in refining sample selection. Extensive experiments demonstrate that our method is superior to the existing state-of-the-art domain adaptation and active learning (AL) methods and comparable with those of full supervision. The code will be released at https://github.com/1061018609/BADA. Gary G. Yen, Chaoqiang Zhao, Qiyu Sun, Wenqi Ren, Lu Sheng, Yang Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | RoMa: Robust Dense Feature MatchingabstractFeature matching is an important computer vision task that involves estimating correspondences between two images of a 3D scene, and dense methods estimate all such correspondences. The aim is to learn a robust model, i.e., a model able to match under challenging real-world changes. In this work, we propose such a model, leveraging frozen pretrained features from the foundation model DINOv2. Al-though these features are significantly more robust than local features trained from scratch, they are inherently coarse. We therefore combine them with specialized ConvNet fine features, creating a precisely localizable feature pyramid. To further improve robustness, we propose a tailored transformer match decoder that predicts anchor probabilities, which enables it to express multimodality. Finally, we propose an improved loss formulation through regression-by-classification with subsequent robust regression. We conduct a comprehensive set of experiments that show that our method, RoMa, achieves significant gains, setting a new state-of-the-art. In particular, we achieve a 36% improvement on the extremely challenging WxBS benchmark. Code is provided at github.com/Parskatt/RoMa. Johan Edstedt, Qiyu Sun, Georg Bökman, Mårten Wadenbäck, Michael Felsberg |
CVPR | 2 |
| 2024 | Demand-Responsive Transport Dynamic Scheduling Optimization Based on Multi-agent Reinforcement Learning Under Mixed Demand
Jianrui Wang, Qiyu Sun, Yang Tang 0001 |
ICANN (4) | 3 |
| 2024 | Temporally Consistent Unpaired Multi-domain Video Translation by Contrastive LearningabstractUnpaired multi-domain video-to-video translation is an attractive solution for diverse video translation, which has to deal with not only unpaired data but also spatio-temporal inconsistency. Most current video-to-video translation models based on cycle consistency introduce optical flow as motion information to achieve spatio-temporal consistency. However, the warping of the optical flow generates a meaningless invisible regions outside the field of view, which is produced by stretching the edge area of the original image and negatively affects the model training. In this work, we propose the Contrastive learning for Multi-domain Video-to-video Translation to replace the cycle consistency, in order to avoid the affects of invisible regions. Specifically, we first introduce synthetic optical flow to maintain spatio-temporal consistency. Then, we use the attention mechanism as the selection principle of positive and negative, and eliminate the features of invisible regions by sorting the feature entropy. Frequency domain information is also used to maintain individual consistency. Experiments on the public datasets Viper and INIT show that our methods is universal across multiple datasets and achieves state-of-the-art performance in generating temporally consistent multi-domain videos. Ruiyang Fan, Qiyu Sun, Ruihao Xia, Yang Tang 0001 |
IJCNN | 2 |
| 2024 | Causal Learning for Heterogeneous Subgroups Based on Nonlinear Causal Kernel ClusteringabstractDue to the challenge posed by multi-source and heterogeneous data collected from diverse environments, causal relationships among features can exhibit variations influenced by different time spans, regions, or strategies. This diversity makes a single causal model inadequate for accurately representing complex causal relationships in all observational data, a crucial consideration in causal learning. To address this challenge, we introduce the nonlinear Causal Kernel Clustering method designed for heterogeneous subgroup causal learning, illuminating variations in causal relationships across diverse subgroups. It comprises two primary components. First, the construction of a sample mapping function forms the basis of the subsequent nonlinear causal kernel. This function assesses the differences in potential nonlinear causal relationships in various samples, supported by our causal identifiability theory. Second, a nonlinear causal kernel is proposed for clustering heterogeneous subgroups. Experimental results showcase the exceptional performance of our method in accurately identifying heterogeneous subgroups and effectively enhancing causal learning, leading to a great reduction in prediction error. Yang Tang 0001, Kexuan Zhang, Qiyu Sun |
IJCNN | 4 |
| 2024 | Scalable Networked Feature Selection with Randomized Algorithm for Robot NavigationabstractWe address the problem of sparse selection of visual features for localizing a team of robots navigating in an unknown environment, where robots can exchange relative position measurements with neighbors. We select a set of the most informative features by anticipating their importance in robots localization by simulating trajectories of robots over a prediction horizon. Through theoretical proofs, we establish a crucial connection between graph Laplacian and the importance of features. We leverage a scalable randomized algorithm for sparse sums of positive semidefinite matrices to efficiently select a set of the most informative features. Vivek Pandey, Arash Amini, Guangyi Liu 0004, Ufuk Topcu, Qiyu Sun, Kostas Daniilidis, Nader Motee |
IROS | 5 |
| 2024 | Learn to Adapt for Self-Supervised Monocular Depth EstimationabstractMonocular depth estimation is one of the fundamental tasks in environmental perception and has achieved tremendous progress by virtue of deep learning. However, the performance of trained models tends to degrade or deteriorate when employed on other new datasets due to the gap between different datasets. Though some methods utilize domain adaptation technologies to jointly train different domains and narrow the gap between them, the trained models cannot generalize to new domains that are not involved in training. To boost the transferability of self-supervised monocular depth estimation models and mitigate the issue of meta-overfitting, we train the model in the pipeline of meta-learning and propose an adversarial depth estimation task. We adopt model-agnostic meta-learning (MAML) to obtain universal initial parameters for further adaptation and train the network in an adversarial manner to extract domain-invariant representations for easing meta-overfitting. In addition, we propose a constraint to impose upon cross-task depth consistency to compel the depth estimation to be identical in different adversarial tasks, which improves the performance of our method and smoothens the training process. Experiments on four new datasets demonstrate that our method adapts quite fast to new domains. Our method trained after 0.5 epoch achieves comparable results with the state-of-the-art methods trained at least 20 epochs. Qiyu Sun, Gary G. Yen, Yang Tang 0001, Chaoqiang Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic SegmentationabstractMost nighttime semantic segmentation studies are based on domain adaptation approaches and image input. However, limited by the low dynamic range of conventional cameras, images fail to capture structural details and boundary information in low-light conditions. Event cameras, as a new form of vision sensors, are complementary to conventional cameras with their high dynamic range. To this end, we propose a novel unsupervised Cross-Modality Domain Adaptation (CMDA) framework to leverage multi-modality (Images and Events) information for nighttime semantic segmentation, with only labels on daytime images. In CMDA, we design the Image Motion-Extractor to extract motion information and the Image Content-Extractor to extract content information from images, in order to bridge the gap between different modalities (Images ⇌ Events) and domains (Day ⇌ Night). Besides, we introduce the first image-event nighttime semantic segmentation dataset. Extensive experiments on both the public image dataset and the proposed image-event dataset demonstrate the effectiveness of our proposed approach. We open-source our code, models, and dataset at https://github.com/XiaRho/CMDA. Ruihao Xia, Chaoqiang Zhao, Meng Zheng 0002, Ziyan Wu 0001, Qiyu Sun, Yang Tang 0001 |
ICCV | 5 |
| 2023 | GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor ScenesabstractThis paper tackles the challenges of self-supervised monocular depth estimation in indoor scenes caused by large rotation between frames and low texture. We ease the learning process by obtaining coarse camera poses from monocular sequences through multi-view geometry to deal with the former. However, we found that limited by the scale ambiguity across different scenes in the training dataset, a naïve introduction of geometric coarse poses cannot play a positive role in performance improvement, which is counter-intuitive. To address this problem, we propose to refine those poses during training through rotation and translation/scale optimization. To soften the effect of the low texture, we combine the global reasoning of vision transformers with an overfitting-aware, iterative self-distillation mechanism, providing more accurate depth guidance coming from the network itself. Experiments on NYUv2, ScanNet, 7scenes, and KITTI datasets support the effectiveness of each component in our framework, which sets a new state-of-the-art for indoor self-supervised monocular depth estimation, as well as outstanding generalization ability. Code and models are available at https://github.com/zxcqlf/GasMono Chaoqiang Zhao, Matteo Poggi, Fabio Tosi, Qiyu Sun, Yang Tang 0001, Stefano Mattoccia |
ICCV | 5 |
| 2023 | Rethinking Unsupervised Domain Adaptation for Nighttime Tracking
Qiyu Sun, Chaoqiang Zhao, Wenqi Ren, Yang Tang 0001 |
ICONIP (14) | 2 |
| 2023 | Molecular Joint Representation Learning via Multi-Modal Information of SMILES and GraphsabstractIn recent years, artificial intelligence has played an important role on accelerating the whole process of drug discovery. Various of molecular representation schemes of different modals (e.g., textual sequence or graph) are developed. By digitally encoding them, different chemical information can be learned through corresponding network structures. Molecular graphs and Simplified Molecular Input Line Entry System (SMILES) are popular means for molecular representation learning in current. Previous works have done attempts by combining both of them to solve the problem of specific information loss in single-modal representation on various tasks. To further fusing such multi-modal imformation, the correspondence between learned chemical feature from different representation should be considered. To realize this, we propose a novel framework of molecular joint representation learning via Multi-Modal information of SMILES and molecular Graphs, called MMSG. We improve the self-attention mechanism by introducing bond-level graph representation as attention bias in Transformer to reinforce feature correspondence between multi-modal information. We further propose a Bidirectional Message Communication Graph Neural Network (BMC GNN) to strengthen the information flow aggregated from graphs for further combination. Numerous experiments on public property prediction datasets have demonstrated the effectiveness of our model. Yang Tang 0001, Qiyu Sun, Luolin Xiong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Perception and Navigation in Autonomous Systems in the Era of Learning: A SurveyabstractAutonomous systems possess the features of inferring their own state, understanding their surroundings, and performing autonomous navigation. With the applications of learning systems, like deep learning and reinforcement learning, the visual-based self-state estimation, environment perception, and navigation capabilities of autonomous systems have been efficiently addressed, and many new learning-based algorithms have surfaced with respect to autonomous visual perception and navigation. In this review, we focus on the applications of learning-based monocular approaches in ego-motion perception, environment perception, and navigation in autonomous systems, which is different from previous reviews that discussed traditional methods. First, we delineate the shortcomings of existing classical visual simultaneous localization and mapping (vSLAM) solutions, which demonstrate the necessity to integrate deep learning techniques. Second, we review the visual-based environmental perception and understanding methods based on deep learning, including deep learning-based monocular depth estimation, monocular ego-motion prediction, image enhancement, object detection, semantic segmentation, and their combinations with traditional vSLAM frameworks. Then, we focus on the visual navigation based on learning systems, mainly including reinforcement learning and deep reinforcement learning. Finally, we examine several challenges and promising directions discussed and concluded in related research of learning systems in the era of computer science and robotics. Yang Tang 0001, Chaoqiang Zhao, Jianrui Wang, Chongzhen Zhang, Qiyu Sun, Wei Xing Zheng 0001, Wenli Du, Feng Qian 0004, Jürgen Kurths |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Distributed algorithms to determine eigenvectors of matrices on spatially distributed networks
Nazar Emirov, Cheng Cheng 0003, Qiyu Sun, Zhihua Qu |
Signal Process. | 3 |
| 2022 | Deep Direct Visual OdometryabstractTraditional monocular direct visual odometry (DVO) is one of the most famous methods to estimate the ego-motion of robots and map environments from images simultaneously. However, DVO heavily relies on high-quality images and accurate initial pose estimation during tracking. With the outstanding performance of deep learning, previous works have shown that deep neural networks can effectively learn 6-DoF (Degree of Freedom) poses between frames from monocular image sequences in the unsupervised manner. However, these unsupervised deep learning-based frameworks cannot accurately generate the full trajectory of a long monocular video because of the scale-inconsistency between each pose. To address this problem, we use several geometric constraints to improve the scale-consistency of the pose network, including improving the previous loss function and proposing a novel scale-to-trajectory constraint for unsupervised training. We call the pose network trained by the proposed novel constraint as TrajNet. In addition, a new DVO architecture, called deep direct sparse odometry (DDSO), is proposed to overcome the drawbacks of the previous direct sparse odometry (DSO) framework by embedding TrajNet. Extensive experiments on the KITTI dataset show that the proposed constraints can effectively improve the scale-consistency of TrajNet when compared with previous unsupervised monocular methods, and integration with TrajNet makes the initialization and tracking of DSO more robust and accurate. Chaoqiang Zhao, Yang Tang 0001, Qiyu Sun, Athanasios V. Vasilakos |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Unsupervised Estimation of Monocular Depth and VO in Dynamic Environments via Hybrid MasksabstractDeep learning-based methods mymargin have achieved remarkable performance in 3-D sensing since they perceive environments in a biologically inspired manner. Nevertheless, the existing approaches trained by monocular sequences are still prone to fail in dynamic environments. In this work, we mitigate the negative influence of dynamic environments on the joint estimation of depth and visual odometry (VO) through hybrid masks. Since both the VO estimation and view reconstruction process in the joint estimation framework is vulnerable to dynamic environments, we propose the cover mask and the filter mask to alleviate the adverse effects, respectively. As the depth and VO estimation are tightly coupled during training, the improved VO estimation promotes depth estimation as well. Besides, a depth-pose consistency loss is proposed to overcome the scale inconsistency between different training samples of monocular sequences. Experimental results show that both our depth prediction and globally consistent VO estimation are state of the art when evaluated on the KITTI benchmark. We evaluate our depth prediction model on the Make3D dataset to prove the transferability of our method as well. Qiyu Sun, Yang Tang 0001, Chongzhen Zhang, Chaoqiang Zhao, Feng Qian 0004, Jürgen Kurths |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Multitask GANs for Semantic Segmentation and Depth Completion With Cycle ConsistencyabstractSemantic segmentation and depth completion are two challenging tasks in scene understanding, and they are widely used in robotics and autonomous driving. Although several studies have been proposed to jointly train these two tasks using some small modifications, such as changing the last layer, the result of one task is not utilized to improve the performance of the other one despite that there are some similarities between these two tasks. In this article, we propose multitask generative adversarial networks (Multitask GANs), which are not only competent in semantic segmentation and depth completion but also improve the accuracy of depth completion through generated semantic images. In addition, we improve the details of generated semantic images based on CycleGAN by introducing multiscale spatial pooling blocks and the structural similarity reconstruction loss. Furthermore, considering the inner consistency between semantic and geometric structures, we develop a semantic-guided smoothness loss to improve depth completion results. Extensive experiments on the Cityscapes data set and the KITTI depth completion benchmark show that the Multitask GANs are capable of achieving competitive performance for both semantic segmentation and depth completion tasks. Chongzhen Zhang, Yang Tang 0001, Chaoqiang Zhao, Qiyu Sun, Zhencheng Ye, Jürgen Kurths |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Masked GAN for Unsupervised Depth and Pose Prediction With Scale ConsistencyabstractPrevious work has shown that adversarial learning can be used for unsupervised monocular depth and visual odometry (VO) estimation, in which the adversarial loss and the geometric image reconstruction loss are utilized as the mainly supervisory signals to train the whole unsupervised framework. However, the performance of the adversarial framework and image reconstruction is usually limited by occlusions and the visual field changes between the frames. This article proposes a masked generative adversarial network (GAN) for unsupervised monocular depth and ego-motion estimations. The MaskNet and Boolean mask scheme are designed in this framework to eliminate the effects of occlusions and impacts of visual field changes on the reconstruction loss and adversarial loss, respectively. Furthermore, we also consider the scale consistency of our pose network by utilizing a new scale-consistency loss, and therefore, our pose network is capable of providing the full camera trajectory over a long monocular sequence. Extensive experiments on the KITTI data set show that each component proposed in this article contributes to the performance, and both our depth and trajectory predictions achieve competitive performance on the KITTI and Make3D data sets. Chaoqiang Zhao, Gary G. Yen, Qiyu Sun, Chongzhen Zhang, Yang Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Preconditioned Gradient Descent Algorithm for Inverse Filtering on Spatially Distributed NetworksabstractGraph filters and their inverses have been widely used in denoising, smoothing, sampling, interpolating and learning. Implementation of an inverse filtering procedure on spatially distributed networks (SDNs) is a remarkable challenge, as each agent on an SDN is equipped with a data processing subsystem with limited capacity and a communication subsystem with confined range due to engineering limitations. In this letter, we introduce a preconditioned gradient descent algorithm to implement the inverse filtering procedure associated with a graph filter having small geodesic-width. The proposed algorithm converges exponentially, and it can be implemented at vertex level and applied to time-varying inverse filtering on SDNs. Cheng Cheng 0003, Nazar Emirov, Qiyu Sun |
IEEE Signal Process. Lett. | 3 |
| 2020 | Design of Nonsubsampled Graph Filter Banks via Lifting SchemesabstractGraph filter banks play a crucial role in the vertex and spectral representation of graph signals. The notion of two-channel nonsubsampled graph filter banks (NSGFBs) on an undirected graph was introduced recently. The absence of downsampling/upsampling operators allows greater flexibility in the design of NSGFBs that achieve perfect reconstruction. However the design of NSGFBs that take the spectral response into account has not been adequately addressed yet. Based on the polynomial/rational lifting scheme, this letter presents a simple method to design NSGFBs with good spectral response and perfect reconstruction. Experimental results will demonstrate the effectiveness of the proposed method in tailoring the spectral responses of the lifted NSGFBs. Application of the NSGFB to denoising will also be considered. Junzheng Jiang, David B. H. Tay, Qiyu Sun, Shan Ouyang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | Phase Retrieval From Multiple-Window Short-Time Fourier MeasurementsabstractIn this paper, we introduce two undirected graphs depending on supports of signals and windows, and we show that the connectivity of those graphs provides either necessary or sufficient conditions to phase retrieval of a signal from magnitude measurements of its multiple-window short-time Fourier transform. Also, we propose an algebraic reconstruction algorithm, and provide an error estimate to our algorithm when magnitude measurements are corrupted by deterministic/random noises. Cheng Cheng 0003, Deguang Han, Qiyu Sun, Guangming Shi |
IEEE Signal Process. Lett. | 4 |
| 2016 | Frequency estimation of sinusoids from nonuniform samples
Alam Abbas Syed, Qiyu Sun, Hassan Foroosh |
Signal Process. | 2 |
| 2015 | Reconstruction of Sparse Wavelet Signals From Partial Fourier MeasurementsabstractIn this letter, we show that high-dimensional sparse wavelet signals can be reconstructed from their partial Fourier measurements on a deterministic sampling set with cardinality about a multiple of signal sparsity. Yang Chen 0047, Cheng Cheng 0003, Qiyu Sun |
IEEE Signal Process. Lett. | 3 |
| 2014 | A Unified Formulation of Gaussian Versus Sparse Stochastic Processes - Part I: Continuous-Domain TheoryabstractWe introduce a general distributional framework that results in a unifying description and characterization of a rich variety of continuous-time stochastic processes. The cornerstone of our approach is an innovation model that is driven by some generalized white noise process, which may be Gaussian or not (e.g., Laplace, impulsive Poisson, or alpha stable). This allows for a conceptual decoupling between the correlation properties of the process, which are imposed by the whitening operator L, and its sparsity pattern, which is determined by the type of noise excitation. The latter is fully specified by a Lévy measure. We show that the range of admissible innovation behavior varies between the purely Gaussian and super-sparse extremes. We prove that the corresponding generalized stochastic processes are well-defined mathematically provided that the (adjoint) inverse of the whitening operator satisfies some Lp bound for p ≥ 1. We present a novel operator-based method that yields an explicit characterization of all Lévy-driven processes that are solutions of constant-coefficient stochastic differential equations. When the underlying system is stable, we recover the family of stationary continuous-time autoregressive moving average processes (CARMA), including the Gaussian ones. The approach remains valid when the system is unstable and leads to the identification of potentially useful generalizations of the Lévy processes, which are sparse and non-stationary. Finally, we show that these processes admit a sparse representation in some matched wavelet domain and provide a full characterization of their transform-domain statistics. Michael Unser, Pouya Dehghani Tafti, Qiyu Sun |
IEEE Trans. Inf. Theory | 3 |
| 2007 | Robust Image Watermarking Based on Multiband Wavelets and Empirical Mode DecompositionabstractIn this paper, we propose a blind image watermarking algorithm based on the multiband wavelet transformation and the empirical mode decomposition. Unlike the watermark algorithms based on the traditional two-band wavelet transform, where the watermark bits are embedded directly on the wavelet coefficients, in the proposed scheme, we embed the watermark bits in the mean trend of some middle-frequency subimages in the wavelet domain. We further select appropriate dilation factor and filters in the multiband wavelet transform to achieve better performance in terms of perceptually invisibility and the robustness of the watermark. The experimental results show that the proposed blind watermarking scheme is robust against JPEG compression, Gaussian noise, salt and pepper noise, median filtering, and ConvFilter attacks. The comparison analysis demonstrate that our scheme has better performance than the watermarking schemes reported recently. Ning Bi, Qiyu Sun, Daren Huang, Zhihua Yang, Jiwu Huang |
IEEE Trans. Image Process. | 2 |