EDBT 2026 Demo / reviewers in the wild / expert
Jun Zhou 0007
dblp:99/3847-7
· DBLP profile ↗
26ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-7753-2238ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Match Any KeypointsabstractPrevious research on sparse feature matching typically involves a staged optimization process of keypoint detection, description, and matching. While it allows the network to adapt to specific inputs, it may limit the network's expressive capability and the overall architectural flexibility. In this study, we rethink the matching framework and propose to directly match any given keypoints, optimizing the matching network in an approximately end-to-end manner. To achieve this, firstly, we dynamically sample random positions within the images as assumed keypoints during training, allowing the network to explore a broader matching space. Secondly, we replace specific descriptors with high-efficiency sparse embeddings at multi levels of the image, facilitating the direct learning of underlying textures. Thirdly, we propose a novel and promising architecture, called Proposal-Guided TRansformer (PGTR), which aggregates context information from neighboring match proposals instead of searching globally with local features. PGTR works especially well under our training approach, and attain a synergistic advantage in terms of performance and efficiency. The overall pipeline achieves outstanding performance on various keypoints without any retraining, and can be flexibly reused when new keypoints emerge, making it valuable for real-world applications. Code will be available. Renjie Pan 0001, Jun Zhou 0007, Hua Yang 0001, Cunyan Li |
IEEE Trans. Image Process. | 3 |
| 2024 | M-RAT: a Multi-grained Retrieval Augmentation Transformer for Image Captioning
Jiayan Song, Renjie Pan 0001, Jun Zhou 0007, Hua Yang 0001 |
ACCV (3) | 3 |
| 2024 | FC-GNN: Recovering Reliable and Accurate Correspondences from InterferencesabstractFinding correspondences between images is essential for many computer vision tasks and sparse matching pipelines have been popular for decades. However, matching noise within and between images, along with inconsistent key-point detection, frequently degrades the matching performance. We review these problems and thus propose: 1) a novel and unified Filtering and Calibrating (FC) approach that jointly rejects outliers and optimizes inliers, and 2) leveraging both the matching context and the underlying image texture to remove matching uncertainties. Under the guidance of the above innovations, we construct Filtering and Calibrating Graph Neural Network (FC-GNN), which follows the FC approach to recover reliable and accurate correspondences from various interferences. FC-GNN conducts an effectively combined inference of contextual and local information through careful embedding and multiple information aggregations, predicting confidence scores and calibration offsets for the input correspondences to jointly filter out outliers and improve pixel-level matching accuracy. Moreover, we exploit the local coherence of matches to perform inference on local graphs, thereby reducing computational complexity. Overall, FC-GNN operates at lightning speed and can greatly boost the performance of diverse matching pipelines across various tasks, showcasing the immense potential of such approaches to become standard and pivotal components of image matching. Code is avaiable at https://github.com/xuy123456/fcgnn. Jun Zhou 0007, Hua Yang 0001, Renjie Pan 0001, Cunyan Li |
CVPR | 2 |
| 2024 | Epicardium Prompt-Guided Real-Time Cardiac Ultrasound Frame-to-Volume Registration
Long Lei, Jun Zhou 0007, Jialun Pei, Baoliang Zhao, Yueming Jin, Jeremy Yuen-Chun Teoh, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 2 |
| 2023 | Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose EstimationabstractOne critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transformer (DFTr) block that can aggregate cross-modality features for improving pose estimation. Unlike existing fusion methods, the proposed DFTr can better model cross-modality semantic correlation by leveraging their semantic similarity, such that globally enhanced features from different modalities can be better integrated for improved information extraction. Moreover, to further improve robustness and efficiency, we introduce a novel weighted vector-wise voting algorithm that employs a non-iterative global optimization strategy for precise 3D keypoint localization while achieving near real-time inference. Extensive experiments show the effectiveness and strong generalization capability of our proposed 3D keypoint voting algorithm. Results on four widely used benchmarks also demonstrate that our method outperforms the state-of-the-art methods by large margins. Code is available at https://github.com/junzastar/DFTr_Voting. Jun Zhou 0007, Kai Chen 0028, Linlin Xu, Qi Dou 0001, Harry Qin |
ICCV | 1 |
| 2023 | Matching-to-Detecting: Establishing Dense and Reliable Correspondences Between Images
Jun Zhou 0007, Renjie Pan 0001, Hua Yang 0001, Cunyan Li |
PRCV (2) | 2 |
| 2022 | How Sound Affects Visual Attention in Omnidirectional VideosabstractIn this paper, we propose a new audio-visual attention dataset that records eye movement for omnidirectional videos with and without sound. We classify the videos into three types according to the number of salient objects and sound sources and analyze the impact of sound on visual attention distribution and inter-observer consistency of viewing area in different types of videos. From the quantitative and qualitative analysis, we find that visual attention will be drawn to and concentrated on the sound source with the presence of sound, especially when there are several visually salient objects and only one sound source. Also, the sound will enhance the consistency of observation areas among viewers to some extent. For more investigations on the impact of sound on visual attention and prospective audio-visual saliency model, we still need further study. Guangtao Zhai, Yucheng Zhu, Jun Zhou 0007, Xiao-Ping Zhang 0002 |
ICIP | 4 |
| 2022 | SO(3)-Pose: SO(3)-Equivariance Learning for 6D Object Pose EstimationabstractAbstract 6D pose estimation of rigid objects from RGB‐D images is crucial for object grasping and manipulation in robotics. Although RGB channels and the depth (D) channel are often complementary, providing respectively the appearance and geometry information, it is still non‐trivial on how to fully benefit from the two cross‐modal data. From the simple yet new observation, when an object rotates, its semantic label is invariant to the pose while its keypoint offset direction is variant to the pose. To this end, we present SO(3)‐Pose, a new representation learning network to explore SO(3)‐equivariant and SO(3)‐invariant features from the depth channel for pose estimation. The SO(3)‐invariant features facilitate to learn more distinctive representations for segmenting objects with similar appearance from RGB channels. The SO(3)‐equivariant features communicate with RGB features to deduce the (missed) geometry for detecting keypoints of an object with the reflective surface from the depth channel. Unlike most of existing pose estimation methods, our SO(3)‐Pose not only implements the information communication between the RGB and depth channels, but also naturally absorbs the SO(3)‐equivariance geometry knowledge from depth images, leading to better appearance and geometry representation learning. Comprehensive experiments show that our method achieves the state‐of‐the‐art performance on three benchmarks. Code is available at https://github.com/phaoran9999/SO3-Pose . Haoran Pan, Jun Zhou 0007, Xuequan Lu, Weiming Wang 0002, Xuefeng Yan 0001, Mingqiang Wei |
Comput. Graph. Forum | 2 |
| 2020 | Viewport-adaptive 360-degree video coding
Qiang Hu 0003, Jun Zhou 0007, Xiaoyun Zhang 0001, Zhiru Shi |
Multim. Tools Appl. | 2 |
| 2018 | Eye Movement Pattern Modeling and Visual Comfort Viewing S3D ImagesabstractStereoscopic-3D (S3D) displays are widely used but present problems related to experiences of visual discomfort for human vision. One aspect of this issue is the movement of the gaze point within different depth fields. Here we aim to analyze the relationship between eye movement patterns and visual comfort experienced when viewing S3D images. Rather than simply labeling eye movement data according to categories such as gaze, saccade and so on, we depoly nonparametric Bayesian method to analyze and cluster several eye movement patterns, and to relate them to visual comfort. The results are relevant to the prediction of visual comfort assessment in S3D images by automatic algorithms. Jun Zhou 0007, Xiao Gu 0001, Shoucheng Zhu, Alan C. Bovik |
VCIP | 2 |
| 2018 | Joint Latent Dirichlet Allocation for Social TagsabstractSocial tags, serving as a textual source of simple but useful semantic metadata to reflect the user preference or describe the web objects, has been widely used in many applications. However, social tags have several unique characteristics, i.e., sparseness and data coupling (i.e., non-IIDness), which makes existing text analysis methods such as LDA not directly applicable. In this paper, we propose a new generative algorithm for social tag analysis named joint latent Dirichlet allocation, which models the generation of tags based on both the users and the objects, and thus accounts for the coupling relationships among social tags. The model introduces two latent factors that jointly influence tag generation: the user's latent interest factor and the object's latent topic factor, formulated as user-topic distribution matrix and object-topic distribution matrix, respectively. A Gibbs sampling approach is adopted to simultaneously infer the above two matrices as well as a topic-word distribution matrix. Experimental results on four social tagging datasets have shown that our model is able to capture more reasonable topics and achieves better performance than five state-of-the-art topic models in terms of the widely used point-wise mutual information metric. In addition, we analyze the learnt topics showing that our model recovers more themes from social tags while LDA may lead the topic vanishing problems, and demonstrate its advantages in the social recommendation by evaluating the retrieval results with mean reciprocal rank metric. Finally, we explore the joint procedure of our model in depth to show the non-IID characteristic of social tagging process. Jiangchao Yao, Yanfeng Wang 0001, Ya Zhang 0002, Jun Sun 0005, Jun Zhou 0007 |
IEEE Trans. Multim. | 5 |
| 2017 | Clothing retrieval with visual attention modelabstractClothing retrieval is a challenging problem in computer vision. With the advance of Convolutional Neural Networks (CNNs), the accuracy of clothing retrieval has been significantly improved. FashionNet [1], a recent study, proposes to employ a set of artificial features in the form of landmarks for clothing retrieval, which are shown to be helpful for retrieval. However, the landmark detection module is trained with strong supervision which requires considerable efforts to obtain. In this paper, we propose a self-learning Visual Attention Model (VAM) to extracts attention maps from clothing images. The VAM is further connected to a global network to form an end-to-end network structure through Impdrop connection which randomly Dropout on the feature maps with the probabilities given by the attention map. Extensive experiments on several widely used benchmark clothing retrieval data sets have demonstrated the promise of the proposed method. We also show that compared to the trivial Product connection, the Impdrop connection makes the network structure more robust when training sets of limited size are used. Yujun Gu, Ya Zhang 0002, Jun Zhou 0007, Xiao Gu 0001 |
VCIP | 4 |
| 2017 | Visual discomfort prediction on stereoscopic 3D images without explicit disparities
Jun Zhou 0007, Jun Sun 0005, Alan C. Bovik |
Signal Process. Image Commun. | 2 |
| 2017 | No-Reference Quality Assessment of Screen Content PicturesabstractRecent years have witnessed a growing number of image and video centric applications on mobile, vehicular, and cloud platforms, involving a wide variety of digital screen content images. Unlike natural scene images captured with modern high fidelity cameras, screen content images are typically composed of fewer colors, simpler shapes, and a larger frequency of thin lines. In this paper, we develop a novel blind/no-reference (NR) model for accessing the perceptual quality of screen content pictures with big data learning. The new model extracts four types of features descriptive of the picture complexity, of screen content statistics, of global brightness quality, and of the sharpness of details. Comparative experiments verify the efficacy of the new model as compared with existing relevant blind picture quality assessment algorithms applied on screen content image databases. A regression module is trained on a considerable number of training samples labeled with objective visual quality predictions delivered by a high-performance full-reference method designed for screen content image quality assessment (IQA). This results in an opinion-unaware NR blind screen content IQA algorithm. Our proposed model delivers computational efficiency and promising performance. The source code of the new model will be available at: https://sites.google.com/site/guke198701/publications. Ke Gu 0001, Jun Zhou 0007, Junfei Qiao 0001, Guangtao Zhai, Weisi Lin, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2016 | Stereoscopic images quality assessment based on deep learningabstractWith the popularity of stereoscopic 3D (S3D) images and videos, many advanced objective quality assessment methods have been proposed to evaluate viewers' Quality of Experience (QoE). Among them, most algorithms take advantages of the disparity maps to extract useful features. On the other hand, deep learning has been one of the hottest research topics during these years, but limited efforts focused on the field in objective quality evaluation of S3D images. In this paper, we propose a S3D image quality assessment (S3D IQA) method based on deep learning. In this method, the Convolutional Restricted Boltzmann Machines (CRBM) combined with Factored Third-Order RBM (FTO-RBM) is considered as learning model to extract feature maps from pre-processed left and right images automatically. Then an improved traversal algorithm based on two pooling strategies is put forward to optimize extracted feature maps, which improves the final quality assessment performance significantly. Experimental results show that our S3D IQA method achieves good performance on 3D databases tested. Kai Wang 0092, Jun Zhou 0007, Xiao Gu 0001 |
VCIP | 2 |
| 2015 | Joint Latent Dirichlet Allocation for non-iid social tagsabstractTopic models have been widely used for analyzing text corpora and achieved great success in applications including content organization and information retrieval. However, different from traditional text data, social tags in the web containers are usually of small amounts, unordered, and non-iid, i.e., it is highly dependent on contextual information such as users and objects. Considering the specific characteristics of social tags, we here introduce a new model named Joint Latent Dirichlet Allocation (JLDA) to capture the relationships among users, objects, and tags. The model assumes that the latent topics of users and those of objects jointly influence the generation of tags. The latent distributions is then inferred with Gibbs sampling. Experiments on two social tag data sets have demonstrated that the model achieves a lower predictive error and generates more reasonable topics. We also present an interesting application of this model to object recommendation. Jiangchao Yao, Ya Zhang 0002, Zhe Xu 0003, Jun Sun 0005, Jun Zhou 0007, Xiao Gu 0001 |
ICME | 5 |
| 2014 | Binocular mismatch induced by luminance discrepancies on stereoscopic imagesabstractLuminance discrepancies between image pairs occur owing to inconsistent parameters between stereoscopic camera devices and from imperfect capture conditions. Such discrepancies induce binocular mismatches and affect the visual comfort that is felt by viewers, as well as their ability to fuse stereoscopic. To better understand and observe this effect, we built a stereoscopic images database of 240 luminance discrepancy images and 30 natural images with subjective scores of visual discomfort and fusion difficulty. Two features, binocular contrast and luminance similarity were extracted to analyze the relationship between the subjective scores and the luminance discrepancies. Structural dissimilarity and average luminance are used to predict the effects of binocular mismatches. The experimental results show that the combination of binocular contrast, structural dissimilarity and average luminance exhibits high consistency with subjective scores of visual discomfort, fusion difficulty and overall binocular mismatches in terms of Spearman's Rank Ordered Correlation Coefficient. Jun Zhou 0007, Jun Sun 0005, Alan C. Bovik |
ICME | 2 |
| 2014 | Query-expanded collaborative representation based classification with class-specific prototypes for object recognition
Jun Zhou 0007, Jun Sun 0005 |
Pattern Recognit. | 2 |
| 2013 | Maximizing Expected Model Change for Active Learning in RegressionabstractActive learning is well-motivated in many supervised learning tasks where unlabeled data may be abundant but labeled examples are expensive to obtain. The goal of active learning is to maximize the performance of a learning model using as few labeled training data as possible, thereby minimizing the cost of data annotation. So far, there is still very limited work on active learning for regression. In this paper, we propose a new active learning framework for regression called Expected Model Change Maximization (EMCM), which aims to choose the examples that lead to the largest change to the current model. The model change is measured as the difference between the current model parameters and the updated parameters after training with the enlarged training set. Inspired by the Stochastic Gradient Descent (SGD) update rule, the change is estimated as the gradient of the loss with respect to a candidate example for active learning. Under this framework, we derive novel active learning algorithms for both linear regression and nonlinear regression to select the most informative examples. Extensive experimental results on the benchmark data sets from UCI machine learning repository have demonstrated that the proposed algorithms are highly effective in choosing the most informative examples and robust to various types of data distributions. Wenbin Cai, Ya Zhang 0002, Jun Zhou 0007 |
ICDM | 3 |
| 2013 | Foreground detection: Combining background subspace learning with object smoothing modelabstractForeground detection is a challenging problem in complex scenes. In this paper, a novel foreground detection method is proposed which combines background subspace learning with object smoothing model. Considering background scenes in consecutive frames are almost the same, they are approximated using an efficient subspace learning technique which is based on 2D images. Due to the pixels of objects are usually clustered, an object smoothing model is adopted where a spatial smoothing constraint is imposed on its values during the estimation, and then it can be solved as a regularized matrix restoration problem with a spatial smoothing constraint. As a result, isolated noises can be suppressed while clustered foreground pixels can be preserved. We test our method on some challenging sequences and compare it with some other techniques. Experimental results show its effectiveness and robustness. Gengjian Xue, Li Song 0001, Jun Sun 0005, Jun Zhou 0007 |
ICME | 4 |
| 2013 | Face Recognition Using Multi-scale ICA Texture Pattern and Farthest Prototype Representation Classification
Jun Zhou 0007, Jun Sun 0005 |
MMM (2) | 2 |
| 2013 | Adaptive high-frequency clipping for improved image quality assessmentabstractIt is widely known that the human visual system (HVS) applies multi-resolution analysis to the scenes we see. In fact, many of the best image quality metrics, e.g. MS-SSIM and IW-PSNR/SSIM are based on multi-scale models. However, in existing multi-scale type of image quality assessment (IQA) methods, the resolution levels are fixed. In this paper, we examine the problem of selecting optimal levels in the multi-resolution analysis to preprocess the image for perceptual quality assessment. According to the contrast sensitivity function (CSF) of the HVS, the sampling of visual information by the human eyes approximates a low-pass process. For images, the amount of information we can extract depends on the size of the image (or the object(s) inside) as well as the viewing distance. Therefore, we proposed a wavelet transform based adaptive high-frequency clipping (AHC) model to approximate the effective visual information that enters the HVS. After the high-frequency clipping, rather than processing separately on each level, we transform the filtered images back to their original resolutions for quality assessment. Extensive experimental results show that on various databases (LIVE, IVC, and Toyama-MICT), performance of existing image quality algorithms (PSNR and SSIM) can be substantially improved by applying the metrics to those AHC model processed images. Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiaokang Yang 0001, Jun Zhou 0007, Wenjun Zhang 0001 |
VCIP | 6 |
| 2012 | Learning a Mahalanobis distance metric via regularized LDA for scene recognitionabstractConstructing a suitable distance metric for scene recognition is a very challenging task due to the huge intra-class variations. In this paper, we propose a novel framework for learning a full parameter matrix in Mahalanobis metric, where the learning process is formulated as a non-negatively constrained minimization problem in a projected space. To fully capture the structure of scenes, we first apply multiple regularized linear discriminant analysis (LDA) to form a candidate projection pool. Second, we adopt the pairwise squared differences of the projected samples as the learning instances. Finally, the diagonal selection matrix is learned through least squares with non-negative L2-norm regularization. Experiments on two datasets in scene recognition show the effectiveness and efficiency of our approach. Jun Zhou 0007, Jun Sun 0005 |
ICIP | 2 |
| 2011 | Adaptive fast DIRECT mode decision algorithm using mode and Lagrangian cost prediction for B frame in H.264/AVCabstractIn this paper, a fast spatial DIRECT mode decision method for B frame in H.264/AVC is proposed. It is based on a statistical analysis on multiple video sequences, and the strong relationship of mode selection and rate-distortion (RD) cost between the current DIRECT macroblock (MB) and the co-located MBs is observed. With the check of mode condition and adaptive threshold of RD cost, the complex mode decision process can be released at an early stage even for small QP cases. Simulation results demonstrate the proposed method can achieve much better performance than the original exhaustive rate-distortion optimization (RDO) based mode decision algorithm by reducing up to 57.1 % of motion estimation (ME) time for IBPBP picture group with only negligible bit increment and quality degradation. Xiaocong Jin, Jun Sun 0005, Jun Zhou 0007, Yiqing Huang 0002, Takeshi Ikenaga |
ICME | 3 |
| 2010 | Texture-based color constancy using local regressionabstractColor constancy endows the machines with the ability of identifying the color regardless of the illuminant. Considering none of single algorithms available is universal, this paper presents a novel combination approach based on local texture features and local regression. To better represent images, local texture features based on integrated Weibull distribution are firstly extracted on the overlapping patches of the images. Then we define a new image distance metric to search for K most similar images of the test image. Incorporating a priori knowledge into the data-driven strategy, we finally combine individual algorithms using a local penalized regression according to the frequency ratio of best single algorithms. Experiment on a widely used dataset shows that the proposed approach outperforms the state-of-the-art single algorithms as well as popular combination approaches. Jun Zhou 0007, Jun Sun 0005, Gengjian Xue |
ICIP | 2 |
| 2007 | Quaternion wavelet phase based stereo matching for uncalibrated images
Jun Zhou 0007, Yi Xu 0001, Xiaokang Yang 0001 |
Pattern Recognit. Lett. | 1 |