Qi Jia 0001

dblp:69/1921-1 · DBLP profile ↗
← Back
44ranked-venue papers
16as first author
30since 2021 · last 2026
0000-0001-5383-0065ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 12 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 12 since 2021
YearPublicationVenuePosition
2026 RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels
abstract
Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their inability to provide precise annotation data for sonar images. Therefore, designing effective object detection methods for sonar images with extremely limited labels is particularly important. To address this, we propose a teacher-student framework called RSOD, which aims to fully learn the characteristics of sonar images and develop a pseudo-label strategy suitable for these images to mitigate the impact of limited labels. First, RSOD calculates a reliability score by assessing the consistency of the teacher's predictions across different views. To leverage this score, we introduce an object mixed pseudo-label method to tackle the shortage of labeled data in sonar images. Finally, we optimize the performance of the student by implementing a reliability-guided adaptive constraint. By taking full advantage of unlabeled data, the student can perform well even in situations with extremely limited labels. Notably, on the UATD dataset, our method, using only 5% of labeled data, achieves results that can compete against those of our baseline algorithm trained on 100% labeled data. We also collected a new dataset to provide more valuable data for research in the field of sonar.
Chengzhou Li, Guanchen Meng, Qi Jia 0001, Jinyuan Liu 0001, Zhu Liu 0004, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
AAAI4
2026 SWG-Fusion: Soft weather-guided multimodal fusion with VLM-assistance for BEV object detection under harsh weather
Weimin Wang 0007, Ruifeng Nie, Yingchi Liu, Long Ma 0002, Chengpei Xu, Qi Jia 0001, Yu Liu 0012, Na Lei
Pattern Recognit.6
2026 Model-aware ellipse detection via parametric correlation learning
abstract
Ellipse detection presents a significant challenge in computer vision and pattern recognition, often hindered by traditional parameter regression methods that fail to account for the unique geometric characteristics and complex parameter interactions of ellipses. These limitations frequently result in imprecise detections, notably with small or partially occluded ellipses. To overcome these challenges, we propose EDNet, a novel ellipse detection network that exploits the geometric properties of ellipses, thus moving beyond the reliance on internal textures. EDNet improves ellipse detection by refining the loss function to better capture the relationship between the error and each parameters during training. It features a LoG-like Edge Detection Module (LEDM) and an Edge Guided Module (EGM) for precise boundary extraction and multi-scale feature enhancement. Additionally, an auxiliary component estimates ellipse vertices, boosting accuracy for occluded ellipses. Experimental results on two wildly-used benchmark datasets demonstrate that EDNet achieves significant improvements, with an average detection accuracy increase of 6% and 10% over leading state-of-the-art models. • We concentrate on the geometric characteristics of ellipse detection via Edge Detection Module and Edge Guided Module. • We design an auxiliary head for the estimation of four ellipse vertices, invoking additional feature attention on these pivotal points. • We establish the relations between the error and geometric characteristics of the ellipse by a model-aware loss function.
Qi Jia 0001, Zezheng Liu, Yu Liu 0012, Yi Wang 0037, Xinwei Xue, Weimin Wang 0007
Signal Process.1
2025 As Pseudo-Label Free as Possible: Leveraging Adaptive Feature Generation for Sparsely Annotated Object Detection
abstract
Compared to fully supervised object detection, training with sparse annotations typically leads to a decline in performance due to insufficient feature diversity. Existing sparsely annotated object detection (SAOD) methods often rely on pseudo-labeling strategies, but these pseudo-labels tend to introduce noise under extreme sparsity. To simultaneously avoid the impact of pseudo-label noise and enhance feature diversity, we propose a novel Adaptive Feature Generation (AdaptFG) model that generates features based on class names. This model integrates a pre-trained CLIP into a VAE-based feature generator, with its core innovation being an Adaptor that adaptively maps CLIP’s semantic embeddings to the object detector domain. Additionally, we introduce inter-class relationship reasoning in detector, which effectively mitigates misclassifications stemming from similar features. Extensive experimental results demonstrate that AdaptFG consistently outperforms state-of-the-art SAOD methods on the PASCAL VOC and MS COCO benchmarks.
Shuilian Yao, Yu Liu 0012, Qi Jia 0001
AAAI3
2025 Physics-Guided Sonar Image Fine-grained Recognition under Scarce Annotations
abstract
Sonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy.
Chengzhou Li, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001
ACM Multimedia3
2025 Bilevel progressive homography estimation via correlative region-focused transformer
Qi Jia 0001, Xiaomei Feng, Wei Zhang 0339, Yu Liu 0012, Nan Pu, Nicu Sebe
Comput. Vis. Image Underst.1
2025 Crossing the Chasm: A practical architecture augmentation for low-quality object detection
Xinwei Xue, Haoze Zheng, Yuechao Gao, Tengyu Ma 0004, Long Ma 0002, Qi Jia 0001
Neurocomputing6
2025 Rectangling for Stitched Image via Pixel-Wise Deformation Learning
abstract
Image rectangling involves filling in the blanks created during image stitching through deformation techniques. However, existing methods still struggle with incomplete filling and distortion of content, ultimately affecting the overall visual impression and potentially hindering subsequent tasks such as recognition. In this work, we design a pixel-wise deformation framework that utilizes explicit edge guidance to maintain consistency of texture and structure, yielding rectangular images with natural structure. Specifically, we decouple motion into region-level and pixel-level components through uniform mesh warping and pixel-wise deformation to precisely rearrange the spatial distribution of all pixels. Uniform deformation preserves local structure within divided patches, while pixel-wise motion coordinates the consistency between patches. Their combination provides robust and accurate pixel-wise offsets for structure-preserved rectangling. To further bolster the consistency of structure and texture, we leverage edge information to establish structural constraints and design an edge-guided enhancement module to aid in restoring fine texture details. Additionally, stitched images encompass both meaningful content and blank spaces, we innovatively incorporate a mask predictor, which acts as a guiding beacon, directing the network's attention solely towards content-rich regions to facilitate precise pixel-wise motion estimation. Experimental results demonstrate that our approach achieves state-of-the-art performance in rectifying irregular boundaries while contributing to downstream visual perception tasks.
Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Weimin Wang 0007, Yuqing Liu 0001, Xinwei Xue
IEEE Trans. Multim.2
2024 Novel Class Discovery for Ultra-Fine-Grained Visual Categorization
abstract
Ultra-fine-grained visual categorization (Ultra-FGVC) aims at distinguishing highly similar sub-categories within fine-grained objects, such as different soybean cultivars. Compared to traditional fine-grained visual categorization, Ultra-FGVC encounters more hurdles due to the small inter-class and large intra-class variation. Given these challenges, relying on human annotation for Ultra-FGVC is impractical. To this end, our work introduces a novel task termed Ultra-Fine-Grained Novel Class Discovery (UFG-NCD), which leverages partially annotated data to identify new categories of unlabeled images for Ultra-FGVC. To tackle this problem, we devise a Region-Aligned Proxy Learning (RAPL) framework, which comprises a Channel-wise Region Alignment (CRA) module and a Semi-Supervised Proxy Learning (SemiPL) strategy. The CRA module is designed to extract and utilize discriminative features from local regions, facilitating knowledge transfer from labeled to unlabeled classes. Furthermore, SemiPL strengthens representation learning and knowledge transfer with proxy-guided supervised learning and proxy-guided contrastive learning. Such techniques leverage class distribution information in the embedding space, improving the mining of subtle differences between labeled and unlabeled ultra-fine-grained classes. Extensive experiments demonstrate that RAPL significantly outperforms baselines across various datasets, indicating its effectiveness in handling the challenges of UFG-NCD. Code is available at https://github.com/SSDUT-Caiyq/UFG-NCD.
Yu Liu 0012, Yaqi Cai, Qi Jia 0001, Binglin Qiu, Weimin Wang 0007, Nan Pu
CVPR3
2024 Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation
abstract
Homography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all planes equally, neglecting the negative impact of regions that differ significantly from the largest approximate planar areas (dominant plane). In this work, we propose a depth-guided dominant plane perception network to achieve unsupervised homography estimation with additional attention on the dominant plane. Specifically, we leverage the depth-wise prior to adaptively detecting the approximate dominant plane, invoking essential scene structures for unsupervised homography estimation. Then, we enhance the corresponding features of the dominant plane and explore their correlations through a specially designed perceptual module. Finally, we employ dominant plane perception on multi-scale features progressively to estimate the homography in a coarse-to-fine manner. Extensive experiments on a large parallax dataset demonstrate that our method improves the alignment performance by 10.29%, yielding more accurate alignment than previous competitive methods.
Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki
ICASSP2
2024 CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning
abstract
Attribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual dependency between the A-O primitive pairs. Inspired by this, we propose a novel A-O disentangled framework for CZSL, namely Class-specified Cascaded Network (CSC-Net). The key insight is to firstly classify one primitive and then specifies the predicted class as a priori for guiding another primitive recognition in a cascaded fashion. To this end, CSCNet constructs Attribute-to-Object and Object-to- Attribute cascaded branches, in addition to a composition branch modeling the two primitives as a whole. Notably, we devise a parametric classifier (ParamCls) to improve the matching between visual and semantic embeddings. By improving the A-O disentanglement, our framework achieves superior results than previous competitive methods.
Yanyi Zhang, Qi Jia 0001, Xin Fan 0001, Yu Liu 0012
ICASSP2
2024 Sketch-Based 3D Shape Retrieval With Multi-View Fusion Transformer
abstract
Sketch-based 3D shape retrieval aims to retrieve similar 3D shapes given a 2D sketch query. Although this task has been studied for years, the inherent cross-modal gap and data imbalance between 2D sketches and 3D shapes remain challenging. To address the problems, we propose a simple and effective framework based on Multi-view Fusion Transformer. To be specific, we project 3D shapes into twelve distinct views, and their CNN features are combined with position embeddings, passing together into a transformer encoder to learn view weights. Then we process them through average pooling and MLPs to obtain the final 3D shape representation. Furthermore, to narrow the data imbalance between 2D sketches and 3D shapes, affine transformation and elastic deformation are fully utilized for sketch augmentation, so as to extract more comprehensive sketch features for feature matching with the multi-view 3D shape representation. Extensive experiments on SHREC13, SHREC14 and PART-SHREC14 datasets demonstrate our method achieves superior performance than previous competitive methods.
Cunjuan Zhu, Dongdong Cui, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012, Michael S. Lew
ICASSP3
2024 Fuzzy Boundary-Guided Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is a challenging task that identifies camouflaged objects from highly similar backgrounds. Existing methods typically treat the whole object equally while neglecting the indistinguishable regions that require more attention than other regions. In this paper, we propose a Fuzzy Boundary-Guided Network (FBG-Net) for camouflaged object detection, which mimics the human behavior that pays more attention to these low-confidence regions when observing objects. Specifically, we devise two main building blocks: (1) Mixed Semantics Aggregation Module (MSAM) to integrate boundary and texture features cumulatively in the high-to-low scales, and (2) Fuzzy Boundary-Guided Module (FBGM) to locate and enhance the low-confidence regions under the guidance of fuzzy boundary. Extensive experiments demonstrate the effectiveness of FBG-Net with superior performance to existing state-of-the-art methods. Code is available at https://github.com/YAOSL98/FBG-Net.
Qi Jia 0001, Shuilian Yao, Youcan Xu, Yu Liu 0012, Dehao Kong, Longin Jan Latecki
ICME1
2024 Joint edge detection learning for recurrent homography estimation
abstract
Homography estimation plays a pivotal role in aligning image pairs across multiple viewpoints. Existing methods focus mainly on texture alignment, whereas overlooking the influence of geometric structures, thereby resulting in inaccurate homography estimation. In this paper, we propose a novel recurrent homography estimation framework with joint edge detection learning. We find that edge detection explores extra anchors for homography estimation, and meanwhile homography provides complementary information of cross views for edge detection refinement. Unlike traditional edge detection applied to individual images, our approach establishes structural consistency constraints to reinforce mutual edges while suppressing unreliable structures. Specifically, the detected edges guide and enhance the texture features through a specifically designed edge-aware fusion module. Ultimately, we recurrently compute the correlation of fusion features from small to large scales for homography regression. Our experimental results demonstrate that the proposed method reduces the matching error by 41.7% than state-of-the-art methods. Furthermore, our network excels in detecting edges with extensive details even under dramatic perspective changes. Code is available at https://github.com/edmandzhao/edge-detection-for-RHE.
Qi Jia 0001, Zikun Zhao, Xiaomei Feng, Jinyuan Liu 0001, Yu Liu 0012, Xinwei Xue
ICME1
2024 Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
abstract
Zero-shot human-object interaction (HOI) detector is capable of generalizing to HOI categories even not encountered during training. Inspired by the impressive zero-shot capabilities offered by CLIP, latest methods strive to leverage CLIP embeddings for improving zero-shot HOI detection. However, these embedding-based methods train the classifier on seen classes only, inevitably resulting in seen-unseen confusion for the model during inference. Besides, we find that using prompt-tuning and adapters further increases the gap between seen and unseen accuracy. To tackle this challenge, we present the first generation-based model using CLIP for zero-shot HOI detection, coined HOIGen. It allows to unlock the potential of CLIP for feature generation instead of feature extraction only. To achieve it, we develop a CLIP-injected feature generator in accordance with the generation of human, object and union features. Then, we extract realistic features of seen samples and mix them with synthetic features together, allowing the model to train seen and unseen classes jointly. To enrich the HOI scores, we construct a generative prototype bank in a pairwise HOI recognition branch, and a multi-knowledge prototype bank in an image-wise HOI recognition branch, respectively. Extensive experiments on HICO-DET benchmark demonstrate our HOIGen achieves superior performance for both seen and unseen classes under various zero-shot settings, compared with other top-performing methods. Code is available at: https://github.com/soberguo/HOIGen
Yu Liu 0012, Weimin Wang 0007, Qi Jia 0001
ACM Multimedia5
2024 Two Teachers Are Better Than One: Semi-supervised Elliptical Object Detection by Dual-Teacher Collaborative Guidance
Yu Liu 0012, Longhan Feng, Qi Jia 0001, Zezheng Liu, Zi-Huang Cao
ACM Multimedia3
2024 PMGNet: Disentanglement and entanglement benefit mutually for compositional zero-shot learning
Yu Liu 0012, Yanyi Zhang, Qi Jia 0001, Weimin Wang 0007, Nan Pu, Nicu Sebe
Comput. Vis. Image Underst.4
2024 WBNet: Weakly-supervised salient object detection via scribble and pseudo-background priors
abstract
Weakly supervised salient object detection (WSOD) methods endeavor to boost sparse labels to get more salient cues in various ways. Among them, an effective approach is using pseudo labels from multiple unsupervised self-learning methods, but inaccurate and inconsistent pseudo labels could ultimately lead to detection performance degradation. To tackle this problem, we develop a new multi-source WSOD framework, WBNet, that can effectively utilize pseudo-background (non-salient region) labels combined with scribble labels to obtain more accurate salient features. We first design a comprehensive salient pseudo-mask generator from multiple self-learning features. Then, we pioneer the exploration of generating salient pseudo-labels via point-prompted and box-prompted Segment-Anything Models (SAM). Then, WBNet leverages a pixel-level Feature Aggregation Module (FAM), a mask-level Transformer-decoder (TFD), and an auxiliary Boundary Prediction Module (EPM) with a hybrid loss function to handle complex saliency detection tasks. Comprehensively evaluated with state-of-the-art methods on five widely used datasets, the proposed method significantly improves saliency detection performance. The code and results are publicly available at https://github.com/yiwangtz/WBNet.
Yi Wang 0037, Ruili Wang 0001, Xiangjian He, Chi Lin 0001, Tianzhu Wang, Qi Jia 0001, Xin Fan 0001
Pattern Recognit.6
2024 Edge-Aware Correlation Learning for Unsupervised Progressive Homography Estimation
abstract
Homography estimation aligns image pairs in cross-views, which is a crucial and fundamental computer vision problem. Existing methods only consider correspondences of texture features for homography estimation, leading to unpleasant artifacts and misalignments introduced by mismatches, especially for low-texture image pairs. In contrast to others, we introduce intuitive structural information as an additional clue that is more sensitive to human vision and low-texture scenarios. In this paper, we propose an edge-aware unsupervised progressive network that couples texture and edge correlation to comprehensively explore potential matching features for homography estimation. To explore robust edge and texture features, we employ a multiscale network to capture feature pyramids with different receptive fields. Then, we design an edge-aware correlation module tailored for homography regression, which plugs in multiscale features to capture accurate correlation maps. Specifically, the edge-aware correlation module leverages the feature-selecting strategy for edge features to capture discriminative matching edges and further guides the texture correlation unit to focus on correctly matched textures. Finally, we leverage multiscale edge-aware correlation maps to predict homography progressively from coarse to fine. Experimental results demonstrate that our proposed method improves PSNR by 11.09% on the real large parallax dataset and reduces matching error by 32.04% on the synthetic COCO dataset, yielding more accurate alignment results than previous state-of-the-art methods.
Xiaomei Feng, Qi Jia 0001, Zikun Zhao, Yu Liu 0012, Xinwei Xue, Xin Fan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Hierarchical Similarity Learning for Aliasing Suppression Image Super-Resolution
abstract
As a highly ill-posed issue, single-image super-resolution (SISR) has been widely investigated in recent years. The main task of SISR is to recover the information loss caused by the degradation procedure. According to the Nyquist sampling theory, the degradation leads to the aliasing effect and makes it hard to restore the correct textures from low-resolution (LR) images. In practice, there are correlations and self-similarities among the adjacent patches in the natural images. This article considers the self-similarity and proposes a hierarchical image super-resolution network (HSRNet) to suppress the influence of aliasing. We consider the SISR issue in the optimization perspective and propose an iterative solution pattern based on the half-quadratic splitting (HQS) method. To explore the texture with local image prior, we design a hierarchical exploration block (HEB) and progressive increase the receptive field. Furthermore, multilevel spatial attention (MSA) is devised to obtain the relations of adjacent feature and enhance the high-frequency information, which acts as a crucial role for visual experience. The experimental result shows that HSRNet achieves better quantitative and visual performance than other works and remits the aliasing more effectively.
Yuqing Liu 0001, Qi Jia 0001, Jian Zhang 0018, Xin Fan 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 A rotation robust shape transformer for cartoon character recognition
Qi Jia 0001, Yi Wang 0037, Xin Fan 0001, Haibin Ling, Longin Jan Latecki
Vis. Comput.1
2023 Image Stitching Based on Multi-Scale Meshes
abstract
Generating high-quality stitching images with a natural structure is a challenging task in computer vision. Recent image stitching methods based on warps failed to suppress the distortion of the images. They often bend the salient lines in the image, which is inconsistent with human perception. In this paper, we succeed in proposing a novelty model called multi-perspective warps for natural image stitching which is related to the density of feature points in the images. With it we can get more precise matching results. Three new energy terms are developed to stitch quality to specify and balance the expected for aligning the vertices of the multi-scale mesh, which can constrain the transformation of the mesh. We also explore and introduce three feature point reconstruction algorithms to enrich the features in the images. Extensive experiments demonstrate that the proposed method outperforms most state-of-the-arts by effectively preserving the linear structure in the image and improving the robustness.
Qi Jia 0001, Nan Pu
ICIP3
2023 Learning Pixel-wise Alignment for Unsupervised Image Stitching
abstract
Image stitching aims to align a pair of images in the same view. Generating precise alignment with natural structures is challenging for image stitching, as there is no wider field-of-view image as a reference, especially in non-coplanar practical scenarios. In this paper, we propose an unsupervised image stitching framework, breaking through the coplanar constraints in homography estimation, yielding accurate pixel-wise alignment under limited overlapping regions. First, we generate a global transformation by an iterative dense feature matching combined with an error control strategy to alleviate the difference introduced by large parallax. Second, we propose a pixel-wise warping network embedded within a large-scale feature extractor and a correlative feature enhancement module to explicitly learn correspondences between the inputs, and generate accurate pixel-level offsets upon novel constraints on both overlapping and non-overlapping regions. Notably, we leverage the pixel-level offsets in the overlapping area to guide the adjustment in the non-overlapping area upon content and structure consistency constraints, rendering a natural transition between two regions and distortions suppression over the entire stitched image. The proposed method achieves state-of-the-art performance that surpasses both traditional and deep learning approaches by a large margin. It also achieves the shortest execution time and has the best generalization ability on the traditional dataset.
Qi Jia 0001, Xiaomei Feng, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki
ACM Multimedia1
2023 Broaden Your Positives: A General Rectification Approach for Novel Class Discovery
Yaqi Cai, Nan Pu, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012
PRCV (4)3
2023 Investigating intrinsic degradation factors by multi-branch aggregation for real-world underwater image enhancement
Xinwei Xue, Long Ma 0002, Qi Jia 0001, Risheng Liu, Xin Fan 0001
Pattern Recognit.4
2023 Characteristic Mapping for Ellipse Detection Acceleration
abstract
It is challenging to characterize the intrinsic geometry of high-degree algebraic curves with lower-degree algebraic curves. The reduction in the curve's degree implies lower computation costs, which is crucial for various practical computer vision systems. In this paper, we develop a characteristic mapping (CM) to recursively degenerate 3n points on a planar curve of n th order to 3(n-1) points on a curve of (n-1) th order. The proposed characteristic mapping enables curve grouping on a line, a curve of the lowest order, that preserves the intrinsic geometric properties of a higher-order curve (ellipse). We prove a necessary condition and derive an efficient arc grouping module that finds valid elliptical arc segments by determining whether the mapped three points are colinear, invoking minimal computation. We embed the module into two latest arc-based ellipse detection methods, which reduces their running time by 25% and 50% on average over five widely used data sets. This yields faster detection than the state-of-the-art algorithms while keeping their precision comparable or even higher. Two CM embedded methods also significantly surpass a deep learning method on all evaluation metrics.
Qi Jia 0001, Xin Fan 0001, Yang Yang 0120, Xuxu Liu, Zhongxuan Luo, Xinchen Zhou, Longin Jan Latecki
IEEE Trans. Image Process.1
2022 Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard Way
abstract
It is challenging to accurately detect camouflaged objects from their highly similar surroundings. Existing methods mainly leverage a single-stage detection fashion, while neglecting small objects with low-resolution fine edges requires more operations than the larger ones. To tackle camouflaged object detection (COD), we are inspired by humans attention coupled with the coarse-to-fine detection strategy, and thereby propose an iterative refinement framework, coined SegMaR, which integrates Segment, Magnify and Reiterate in a multi-stage detection fashion. Specifically, we design a new discriminative mask which makes the model attend on the fixation and edge regions. In addition, we leverage an attention-based sampler to magnify the object region progressively with no need of enlarging the image size. Extensive experiments show our SegMaR achieves remarkable and consistent improvements over other state-of-the-art methods. Especially, we surpass two competitive methods 7.4% and 20.0% respectively in average over standard evaluation metrics on small camouflaged objects. Additional studies provide more promising insights into Seg-MaR, including its effectiveness on the discriminative mask and its generalization to other network architectures. Code is available at https://github.com/dlut-dimt/SegMaR.
Qi Jia 0001, Shuilian Yao, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Zhongxuan Luo
CVPR1
2022 Cross-SRN: Structure-Preserving Super-Resolution Network With Cross Convolution
abstract
It is challenging to restore low-resolution (LR) images to super-resolution (SR) images with correct and clear details. Existing deep learning works almost neglect the inherent structural information of images, which acts as an important role for visual perception of SR results. In this paper, we design a hierarchical feature exploitation network to probe and preserve structural information in a multi-scale feature fusion manner. First, we propose a cross convolution upon traditional edge detectors to localize and represent edge features. Then, cross convolution blocks (CCBs) are designed with feature normalization and channel attention to consider the inherent correlations of features. Finally, we leverage multi-scale feature fusion group (MFFG) to embed the cross convolution blocks and develop the relations of structural features in different scales hierarchically, invoking a lightweight structure-preserving network named as Cross-SRN. Experimental results demonstrate the Cross-SRN achieves competitive or superior restoration performances against the state-of-the-art methods with accurate and clear structural details. Moreover, we set a criterion to select images with rich structural textures. The proposed Cross-SRN outperforms the state-of-the-art methods on the selected benchmark, which demonstrates that our network has a significant advantage in preserving edges.
Yuqing Liu 0001, Qi Jia 0001, Xin Fan 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Leveraging Line-Point Consistence To Preserve Structures for Wide Parallax Image Stitching
abstract
Generating high-quality stitched images with natural structures is a challenging task in computer vision. In this paper, we succeed in preserving both local and global geometric structures for wide parallax images, while reducing artifacts and distortions. A projective invariant, Characteristic Number, is used to match co-planar local sub-regions for input images. The homography between these well-matched sub-regions produces consistent line and point pairs, suppressing artifacts in overlapping areas. We explore and introduce global collinear structures into an objective function to specify and balance the desired characters for image warping, which can preserve both local and global structures while alleviating distortions. We also develop comprehensive measures for stitching quality to quantify the collinearity of points and the discrepancy of matched line pairs by considering the sensitivity to linear structures for human vision. Extensive experiments demonstrate the superior performance of the proposed method over the state-of-the-art by presenting sharp textures and preserving prominent natural structures in stitched images. Especially, our method not only exhibits lower errors but also the least divergence across all test images. Code is available at https://github.com/dut-media-lab/Image-Stitching.
Qi Jia 0001, Zhengjun Li, Xin Fan 0001, Shiyu Teng, Xinchen Ye, Longin Jan Latecki
CVPR1
2021 Underwater Species Detection using Channel Sharpening Attention
abstract
With the continuous exploration of marine resources, underwater artificial intelligent robots play an increasingly important role in the fish industry. However, the detection of underwater objects is a very challenging problem due to the irregular movement of underwater objects, the occlusion of sand and rocks, the diversity of water illumination, and the poor visibility and low color contrast in the underwater environment. In this article, we first propose a real-world underwater object detection dataset (UODD), which covers more than 3K images of the most common aquatic products. Then we propose Channel Sharpening Attention Module (CSAM) as a plug-and-play module to further fuse high-level image information, providing the network with the privilege of selecting feature maps. Fusion of original images through CSAM can improve the accuracy of detecting small and medium objects, thereby improving the overall detection accuracy. We also use Water-Net as a preprocessing method to remove the haze and color cast in complex underwater scenes, which shows a satisfactory detection result on small-sized objects. In addition, we use the class weighted loss as the training loss, which can accurately describe the relationship between classification and precision of bounding boxes of targets, and the loss function converges faster during the training process. Experimental results show that the proposed method reaches a maximum AP of 50.1%, outperforming other traditional and state-of-the-art detectors. In addition, our model only needs an average inference time of 25.4 ms per image, which is quite fast and might suit the real-time scenario.
Lihao Jiang, Yi Wang 0037, Qi Jia 0001, Shengwei Xu, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Xinwei Xue, Ruili Wang 0001
ACM Multimedia3
2020 An Efficient Ellipse Detector Based On Region Detection And Arc Pruning
abstract
Detecting ellipses accurately and efficiently for real-world images is crucial for various visual-based applications. Most existing methods employ detection strategies throughout images, while most time is spent on the non-ellipse region. Meanwhile, small ellipses are often miss-detected due to the low resolution and fixed parameters of detectors. In this paper, we proposed an effective ellipse detector benefiting from the region detection method, which provides a basic estimation on the region and size of ellipses. Then, a two-level arc pruning strategy is proposed to detect ellipses efficiently while limiting false-positive and false-negative results. Furthermore, for the pre-estimated region without detected ellipses, interpolation method is employed to enlarge the target region, which makes small and blur ellipses to be detected. Experimental results demonstrate that the proposed method achieves competitive accuracy compared with the state-of-the-art methods.
Ruike Zhang, Jingchao Liang, Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo
ICIP4
2020 Multi-Scale Features Joint Rain Removal For Single Image
abstract
The presence of rain and haze often cause degradation of images. Therefore, it is important to remove rain or haze and recover the background in outdoor vision systems. Due to the limited size of the network acceptance domain, the pixel value of each spatial position can only be inferred from the surrounding small local area; thus, it is often difficult to remove long rain streaks using existing methods. Therefore, we propose a feature joint dense network (FJDN) to extract multi-scale aggregation features. First, we design a multiscale feature extraction module that uses four dilated convolutional layers to extract multi-scale features. These multi-scale features are then combined into one feature map. We also aggregate three multi-scale features in feature joint dense block (FJDB). By using multi-scale features, we can effectively detect rain streaks of different lengths. Finally, we perform multiple experiments to visually and quantitatively compare our method with several existing methods, demonstrating its superiority. The proposed method is also applied to image dehazing.
Xinwei Xue, Zhenhua Hao, Ying Ding 0006, Qi Jia 0001, Risheng Liu
ICIP4
2020 Coupling Deep Textural and Shape Features for Sketch Recognition
abstract
Recognizing freehand sketches with high arbitrariness is such a great challenge that the automatic recognition rate has reached a ceiling in recent years. In this paper, we explicitly explore the shape properties of sketches, which has almost been neglected before in the context of deep learning, and propose a sequential dual learning strategy that combines both shape and texture features. We devise a two-stage recurrent neural network to balance these two types of features. Our architecture also considers stroke orders of sketches to reduce the intra-class variations of input features. Extensive experiments on the TU-Berlin benchmark set show that our method achieves over 90% recognition rate for the first time on this task, outperforming both humans and state-of-the-art algorithms by over 19 and 7.5 percentage points, respectively. Especially, our approach can distinguish the sketches with similar textures but different shapes more effectively than recent deep networks. Based on the proposed method, we develop an on-line sketch retrieval and imitation application to teach children or adults to draw. The application is available as Sketch.Draw.
Qi Jia 0001, Xin Fan 0001, Meiyu Yu, Yuqing Liu 0001, Dingrong Wang, Longin Jan Latecki
ACM Multimedia1
2018 Line matching based on line-points invariant and local homography
Qi Jia 0001, Xin Fan 0001, Xinkai Gao, Meiyu Yu, Zhongxuan Luo
Pattern Recognit.1
2017 Efficiently building 3D line model with points
abstract
3D modeling is a popular topic in the field of computer vision. Most of the existing methods based on the interest point matches or the line matches. The point-based methods are more mature, but usually less intuitive. Lines can provide more structural information and thus more intuitive, however, lines have more complex geometric properties, which makes the line-based methods usually time consuming. In this paper, we proposed a simple and effective approach of 3D modeling, which takes the intersections of the real lines and the virtual lines passing through interest points to construct the 3D models, and use a simple invariant to filter out wrong matches. Thus we can use the mature technology of point-based methods while the 3D point cloud maintain the structural information of lines. Experiments show our method can build meaningful 3D point cloud efficiently.
Xinkai Gao, Qi Jia 0001, He Guo 0001
ICIP2
2017 A Fast Ellipse Detector Using Projective Invariant Pruning
abstract
Detecting elliptical objects from an image is a central task in robot navigation and industrial diagnosis, where the detection time is always a critical issue. Existing methods are hardly applicable to these real-time scenarios of limited hardware resource due to the huge number of fragment candidates (edges or arcs) for fitting ellipse equations. In this paper, we present a fast algorithm detecting ellipses with high accuracy. The algorithm leverages a newly developed projective invariant to significantly prune the undesired candidates and to pick out elliptical ones. The invariant is able to reflect the intrinsic geometry of a planar curve, giving the value of -1 on any three collinear points and +1 for any six points on an ellipse. Thus, we apply the pruning and picking by simply comparing these binary values. Moreover, the calculation of the invariant only involves the determinant of a 3×3 matrix. Extensive experiments on three challenging data sets with 648 images demonstrate that our detector runs 20%-50% faster than the state-of-the-art algorithms with the comparable or higher precision.
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Lianbo Song, Tie Qiu 0001
IEEE Trans. Image Process.1
2016 Novel Coplanar Line-Points Invariants for Robust Line Matching Across Views
Qi Jia 0001, Xinkai Gao, Xin Fan 0001, Zhongxuan Luo, Ziyao Chen
ECCV (8)1
2016 Cross-view action matching using a novel projective invariant on non-coplanar space-time points
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Kang Huyan, Zezhou Li
Multim. Tools Appl.1
2016 Hierarchical projective invariant contexts for shape recognition
Qi Jia 0001, Xin Fan 0001, Yu Liu 0012, Zhongxuan Luo, He Guo 0001
Pattern Recognit.1
2016 3D facial landmark localization using texture regression via conformal mapping
Xin Fan 0001, Qi Jia 0001, Kang Huyan, Xianfeng Gu, Zhongxuan Luo
Pattern Recognit. Lett.2
2015 Online gesture-based interaction with visual oriental characters based on manifold learning
Yi Wang 0037, Xin Fan 0001, Xiangjian He, Qi Jia 0001, Renjie Gao
Signal Process.5
2014 A new geometric descriptor for symbols with affine deformations
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Yu Liu 0012, He Guo 0001
Pattern Recognit. Lett.1
2013 A shape matching framework using metric partition constraint
abstract
The crucial problem for shape matching is to balance between discrimination power and computation complexity. Popular solutions mainly rely on either global or local information of shape contours, and neglect their intrinsic correlation. But the methods that combine both information may bring high computation complexity. In this paper, we present a shape matching framework, in which a novel shape descriptor named metric partition constraint (MPC) is proposed, and many metric methods can be included. The metric information is used to bridge the local points and the global shape. Meanwhile, we devise a partition smoothing process to improve the robustness to local deformation. Finally, Comprehensive comparisons with the classical shape context and other latest methods on standard datasets show the excellent performance in terms of precision while retaining computational efficiency.
Yu Liu 0012, Qi Jia 0001, He Guo 0001, Xin Fan 0001
ICIP2
2013 A shape descriptor based on new projective invariants
abstract
Great attention has been devoted to the development of shape descriptors that is the key to object recognition. Previous works have great success on either relatively simple shapes or limited transformations, e.g., translation, rotation and scaling. We propose a new projective invariant, named characteristic number (CN) that includes more points for complex shapes with rich inner structures. Moreover, we build a novel shape descriptor with CN values calculated on triangles that cover the convex hull of a shape. The matching based on the descriptor also runs fast since only one initial point for the triangular coverage needs to align based on its CN value prior to the matching. The performance of the proposed descriptor is validated by the experiments compared with the classical shape context (SC) and recently developed cross ratio spectrum (CRS) on 32 logos of television networks with a wide range of transformations (512 images in total).
Zhongxuan Luo, Daiyun Luo, Xin Fan 0001, Xinchen Zhou, Qi Jia 0001
ICIP5