Qingxuan Lv

dblp:260/2895 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-from-gradients
abstract
Surface-from-gradients (SfG) aims to recover a three-dimensional (3D) surface from its gradients. Traditional methods encounter significant challenges in achieving high accuracy and handling high-resolution inputs, particularly facing the complex nature of discontinuities and the inefficiencies associated with large-scale linear solvers. Although recent advances in deep learning, such as photometric stereo, have enhanced normal estimation accuracy, they do not fully address the intricacies of gradient-based surface reconstruction. To overcome these limitations, we propose a Fourier neural operator-based Numerical Integration Network (FNIN) within a two-stage optimization framework. In the first stage, our approach employs an iterative architecture for numerical integration, harnessing an advanced Fourier neural operator to approximate the solution operator in Fourier space. Additionally, a self-learning attention mechanism is incorporated to effectively detect and handle discontinuities. In the second stage, we refine the surface reconstruction by formulating a weighted least squares problem, addressing the identified discontinuities rationally. Extensive experiments demonstrate that our method achieves significant improvements in both accuracy and efficiency compared to current state-of-the-art solvers. This is particularly evident in handling high-resolution images with complex data, achieving errors of fewer than 0.1 mm on tested objects.
Jiaqi Leng 0002, Yakun Ju, Yuanxu Duan, Jiangnan Zhang, Qingxuan Lv, Zuxuan Wu, Hao Fan 0004
AAAI5
2025 UWStereo: A Large Synthetic Dataset for Underwater Stereo Matching
abstract
Despite recent advances in stereo matching, the extension to intricate underwater settings remains unexplored, primarily owing to: 1) the reduced visibility, low contrast, and other adverse effects of underwater images; 2) the difficulty in obtaining ground truth data for training deep learning models, i.e. simultaneously capturing an image and estimating its corresponding pixel-wise depth information in underwater environments. To enable further advance in underwater stereo matching, we introduce a large synthetic dataset called UWStereo. Our dataset includes 29,568 synthetic stereo image pairs with dense and accurate disparity annotations for left view. We design four distinct underwater scenes filled with diverse objects such as corals, ships and robots. We also induce additional variations in camera model, lighting, and environmental effects. In comparison with existing underwater datasets, UWStereo is superior in terms of scale, variation, annotation, and photo-realistic image quality. To substantiate the efficacy of the UWStereo dataset, we undertake a comprehensive evaluation compared with eleven state-of-the-art algorithms as benchmarks. The results indicate that current models still struggle to generalize to new domains. Hence, we design a new strategy that learns to reconstruct cross domain masked images before stereo matching training and integrate a cross view attention enhancement module that aggregates long-range content information to enhance the generalization ability.
Qingxuan Lv, Junyu Dong, Yuezun Li, Sheng Chen 0001, Hui Yu 0001, Shu Zhang 0002, Wenhan Wang
IEEE Trans. Circuits Syst. Video Technol.1
2025 PhyTracker: An Online Tracker for Phytoplankton
abstract
Phytoplankton, a crucial component of aquatic ecosystems, requires efficient monitoring to understand marine ecological processes and environmental conditions. Traditional phytoplankton monitoring methods, relying on non-in situ observations, are time-consuming and resource-intensive, limiting timely analysis. To address these limitations, we introduce PhyTracker, an intelligent in situ tracking framework designed for automatic tracking of phytoplankton. PhyTracker overcomes significant challenges unique to phytoplankton monitoring, such as constrained mobility within water flow, inconspicuous appearance, and the presence of impurities. Our method incorporates three innovative modules: a Texture-enhanced Feature Extraction (TFE) module, an Attention-enhanced Temporal Association (ATA) module, and a Flow-agnostic Movement Refinement (FMR) module. These modules enhance feature capture, differentiate between phytoplankton and impurities, and refine movement characteristics, respectively. Extensive experiments on the PMOT dataset validate the superiority of PhyTracker in phytoplankton tracking, and additional tests on the MOT dataset demonstrate its general applicability, outperforming conventional tracking methods. This work highlights key differences between phytoplankton and traditional objects, offering an effective solution for phytoplankton monitoring.
Qingxuan Lv, Yuezun Li, Zhiqiang Wei 0002, Junyu Dong
IEEE Trans. Circuits Syst. Video Technol.2
2025 HG-SFDA: HyperGraph Learning Meets Source-Free Unsupervised Domain Adaptation
abstract
Source-Free unsupervised Domain Adaptation (SFDA) aims to classify target samples by only accessing a pre-trained source model and unlabelled target samples. Since no source data is available, transferring the knowledge from the source domain to the target domain is challenging. Existing methods normally exploit the pair-wise relation among target samples and attempt to discover their correlations by clustering these samples based on semantic features. The drawbacks of these methods include: 1) the pair-wise relation is limited to exposing the underlying correlations of two more samples, hindering the exploration of the structural information embedded in the target domain; and 2) the clustering process only relies on the semantic feature, while overlooking the critical effect of domain shift, i.e., the distribution differences between the source and target domains. To address these issues, we propose a new SFDA method that exploits the high-order neighborhood relation and explicitly takes the domain shift effect into account. Specifically, we formulate the SFDA as a hypergraph learning problem and construct hyperedges to explore the deep structural and context information among multiple samples. Moreover, we integrate a self-loop strategy into the constructed hypergraph to elegantly introduce the domain uncertainty of each sample. By clustering these samples based on hyperedges, both the semantic feature and domain shift effects are considered. We then describe an adaptive relation-based objective to tune the model with soft attention levels for all samples. Extensive experiments are conducted on Office-31, Office-Home, VisDA, DomainNet-126 and PointDA-10 datasets. The results demonstrate the superiority of our method over state-of-the-art counterparts. Our code is avaliable at https://github.com/OUC-POVA/HG-SFDA.
Jinkun Jiang, Qingxuan Lv, Yuezun Li, Yong Du 0003, Junyu Dong, Sheng Chen 0001, Hui Yu 0001
IEEE Trans. Image Process.2
2024 Multiview adaptive attention pooling for image-text retrieval
Yunlai Ding, Jiaao Yu 0001, Qingxuan Lv, Junyu Dong, Yuezun Li
Knowl. Based Syst.3
2024 DomainForensics: Exposing Face Forgery Across Domains via Bi-Directional Adaptation
abstract
Recent DeepFake detection methods have shown excellent performance on public datasets but are significantly degraded on new forgeries. Solving this problem is important, as new forgeries emerge daily with the continuously evolving generative techniques. Many efforts have been made for this issue by seeking the commonly existing traces empirically on data level. In this paper, we rethink this problem and propose a new solution from the unsupervised domain adaptation perspective. Our solution, called DomainForensics, aims to transfer the forgery knowledge from known forgeries (fully labeled source domain) to new forgeries (label-free target domain). Unlike recent efforts, our solution does not focus on data view but on learning strategies of DeepFake detectors to capture the knowledge of new forgeries through the alignment of domain discrepancies. In particular, unlike the general domain adaptation methods which consider the knowledge transfer in the semantic class category, thus having limited application, our approach captures the subtle forgery traces. We describe a new bi-directional adaptation strategy dedicated to capturing the forgery knowledge across domains. Specifically, our strategy considers both forward and backward adaptation, to transfer the forgery knowledge from the source domain to the target domain in forward adaptation and then reverse the adaptation from the target domain to the source domain in backward adaptation. In forward adaptation, we perform supervised training for the DeepFake detector in the source domain and jointly employ adversarial feature adaptation to transfer the ability to detect manipulated faces from known forgeries to new forgeries. In backward adaptation, we further improve the knowledge transfer by coupling adversarial adaptation with self-distillation on new forgeries. This enables the detector to expose new forgery features from unlabeled data and avoid forgetting the known knowledge of known forgery. Extensive experiments demonstrate that our method is surprisingly effective in exposing new forgeries, and can be plug-and-play on other DeepFake detection architectures.
Qingxuan Lv, Yuezun Li, Junyu Dong, Sheng Chen 0001, Hui Yu 0001, Huiyu Zhou 0001, Shu Zhang 0002
IEEE Trans. Inf. Forensics Secur.1
2024 MobileSky: Real-Time Sky Replacement for Mobile AR
abstract
We present MobileSky, the first automatic method for real-time high-quality sky replacement for mobile AR applications. The primary challenge of this task is how to extract sky regions in camera feed both quickly and accurately. While the problem of sky replacement is not new, previous methods mainly concern extraction quality rather than efficiency, limiting their application to our task. We aim to provide higher quality, both spatially and temporally consistent sky mask maps for all camera frames in real time. To this end, we develop a novel framework that combines a new deep semantic network called FSNet with novel post-processing refinement steps. By leveraging IMU data, we also propose new sky-aware constraints such as temporal consistency, position consistency, and color consistency to help refine the weakly classified part of the segmentation output. Experiments show that our method achieves an average of around 30 FPS on off-the-shelf smartphones and outperforms the state-of-the-art sky replacement methods in terms of execution speed and quality. In the meantime, our mask maps appear to be visually more stable across frames. Our fast sky replacement method enables several applications, such as AR advertising, art making, generating fantasy celestial objects, visually learning about weather phenomena, and advanced video-based visual effects. To facilitate future research, we also create a new video dataset containing annotated sky regions with IMU data.
Xinjie Wang 0003, Qingxuan Lv, Jing Zhang 0038, Zhiqiang Wei 0002, Junyu Dong, Hongbo Fu 0001, Zhipeng Zhu, Xiaogang Jin 0001
IEEE Trans. Vis. Comput. Graph.2
2023 Energy Transfer Contrast Network for Unsupervised Domain Adaption
Jiajun Ouyang, Qingxuan Lv, Shu Zhang 0002, Junyu Dong
MMM (2)2
2023 LaFea: Learning Latent Representation Beyond Feature for Universal Domain Adaptation
abstract
Universal Domain Adaptation (UniDA) is a recent advent problem that aims to transfer the knowledge from the source domain to the target domain without any prior knowledge on label sets. The main challenge is to separate common samples from private samples in the target domain. In general, existing methods achieve this goal by performing domain adaptation only on the features extracted by the backbone networks. However, solely relying on the learning of the backbone network may not fully exploit the effectiveness of features, due to that 1) the discrepancy between two domains can naturally distract the learning of backbone network and 2) the irrelevant content of samples (e. g., backgrounds) likely goes through the backbone network, and accordingly may hinder the learning of domain-informative features. To this end, we describe a new method to provide extra guidance to the learning of the backbone network based on the latent representation beyond features (LaFea). We are motivated by the fact that the latent representation can be learned to contain the domain-relevant information scattered in features, and the learning of this latent representation can naturally promote the effectiveness of corresponding features in return. To achieve this goal, we develop a simple GAN-style architecture to transform features into the latent representation and propose new objectives to adversarially learn this representation. It should be noted that the latent representation only serves as an auxiliary in training, but it is not needed in inference. Extensive experiments on four datasets corroborate the superiority of our method compared to the state-of-the-arts.
Qingxuan Lv, Yuezun Li, Junyu Dong, Ziqian Guo
IEEE Trans. Circuits Syst. Video Technol.1
2022 Parallel Complement Network for Real-Time Semantic Segmentation of Road Scenes
abstract
Real-time semantic segmentation is in intense demand for the application of autonomous driving. Most of the semantic segmentation models tend to use large feature maps and complex structures to enhance the representation power for high accuracy. However, these inefficient designs increase the amount of computational costs, which hinders the model to be applied on autonomous driving. In this paper, we propose a lightweight real-time segmentation model, named Parallel Complement Network (PCNet), to address the challenging task with fewer parameters. A Parallel Complement layer is introduced to generate complementary features with a large receptive field. It provides the ability to overcome the problem of similar feature encoding among different classes, and further produces discriminative representations. With the inverted residual structure, we design a Parallel Complement block to construct the proposed PCNet. Extensive experiments are carried out on challenging road scene datasets, i.e., CityScapes and CamVid, to make comparison against several state-of-the-art real-time segmentation models. The results show that our model has promising performance. Specifically, PCNet* achieves 72.9% Mean IoU on CityScapes using only 1.5M parameters and reaches 79.1 FPS with$1024\times 2048$resolution images on GTX 2080Ti. Moreover, our proposed system achieves the best accuracy when being trained from scratch.
Qingxuan Lv, Xin Sun 0003, Changrui Chen, Junyu Dong, Huiyu Zhou 0001
IEEE Trans. Intell. Transp. Syst.1
2021 GPNet: Gated pyramid network for semantic segmentation
Yu Zhang 0165, Xin Sun 0003, Junyu Dong, Changrui Chen, Qingxuan Lv
Pattern Recognit.5