VLDB 2026 Research / reviewers in the wild / expert
Yongjie Shi
dblp:195/9114
· DBLP profile ↗
23ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic multi-view graph network for multi-parameter condition monitoring of HVDC converter transformers
Yongjie Shi, Tongqiang Yi, Jiale Tian |
Neurocomputing | 1 |
| 2026 | A Novel Method for Robust Fault Diagnosis in Hydropower Units Using Multi-Sensor Fusion and Graph Hierarchical Clustering
Tongqiang Yi, Yongjie Shi, Wenyang Lei |
Knowl. Based Syst. | 2 |
| 2024 | Tensor decompositions for temporal knowledge graph completion with time perspective
Jinfa Yang, Xianghua Ying, Yongjie Shi, Bowei Xing |
Expert Syst. Appl. | 3 |
| 2024 | Improving static and temporal knowledge graph embedding using affine transformations of entitiesabstractTo find a suitable embedding for a knowledge graph (KG) remains a big challenge nowadays. By measuring the distance or plausibility of triples and quadruples in static and temporal knowledge graphs, many reliable knowledge graph embedding (KGE) models are proposed. However, these classical models may not be able to represent and infer various relation patterns well, such as TransE cannot represent symmetric relations, DistMult cannot represent inverse relations, RotatE cannot represent multiple relations, etc.. In this paper, we improve the ability of these models to represent various relation patterns by introducing the affine transformation framework. Specifically, we first utilize a set of affine transformations related to each relation or timestamp to operate on entity vectors, and then these transformed vectors can be applied not only to static KGE models, but also to temporal KGE models. The main advantage of using affine transformations is their good geometry properties with interpretability. Our experimental results demonstrate that the proposed intuitive design with affine transformations provides a statistically significant increase in performance with adding a few extra processing steps and keeping the same number of embedding parameters. Taking TransE as an example, we employ the scale transformation (the special case of an affine transformation). Surprisingly, it even outperforms RotatE to some extent on various datasets. We also introduce affine transformations into RotatE, Distmult, ComplEx, TTransE and TComplEx respectively, and experiments demonstrate that affine transformations consistently and significantly improve the performance of state-of-the-art KGE models on both static and temporal knowledge graph benchmarks. Jinfa Yang, Xianghua Ying, Yongjie Shi, Ruibin Wang |
J. Web Semant. | 3 |
| 2023 | Improving point cloud classification and segmentation via parametric veronese mappingabstractDeep learning based 3D point cloud classification and segmentation has achieved remarkable success. Existing methods are usually implemented in the original space with 3D coordinates as inputs. However, we find that point networks taking only information of first-order coordinates hardly learn geometric features of higher order, such as point cloud normals or poses. In this study, we propose to map the input point clouds into a non-linear space to facilitate networks learning and leveraging high-order features. Firstly, we design the Parametric Veronese Mapping (PVM) function which automatically learns to map point clouds into a non-linear space. As a result, the mapped point clouds are enriched with high-order elements and maintain the basic point set properties as in the original 3D space. We can then exploit existing networks to learn high-order features from mapped point clouds. Secondly, we contribute a two-stage transformation learning module that modifies the previous one-stage module to better leverage high-order features for aligning point clouds in the projective space. Finally, an interaction module is designed to learn more discriminative features by aggregating information from both the original and projective space. Extensive experiments demonstrate that our method successfully improves the ability of most existing networks to learn high-order features and thus contributing to more accurate classification and segmentation. Moreover, the resulting models show stronger robustness to affine transformations and real-world perturbations. Ruibin Wang, Xianghua Ying, Bowei Xing, Xin Tong 0007, Taiyan Chen, Jinfa Yang, Yongjie Shi |
Pattern Recognit. | 7 |
| 2022 | Learning Hierarchy-Aware Quaternion Knowledge Graph Embeddings with Representing Relations as 3D RotationsabstractKnowledge graph embedding aims to represent entities and relations as low-dimensional vectors, which is an effective way for predicting missing links. It is crucial for knowledge graph embedding models to model and infer various relation patterns, such as symmetry/antisymmetry. However, many existing approaches fail to model semantic hierarchies, which are common in the real world. We propose a new model called HRQE, which represents entities as pure quaternions. The relational embedding consists of two parts: (a) Using unit quaternions to represent the rotation part in 3D space, where the head entities are rotated by the corresponding relations through Hamilton product. (b) Using scale parameters to constrain the modulus of entities to make them have hierarchical distributions. To the best of our knowledge, HRQE is the first model that can encode symmetry/antisymmetry, inversion, composition, multiple relation patterns and learn semantic hierarchies simultaneously. Experimental results demonstrate the effectiveness of HRQE against some of the SOTA methods on four well-established knowledge graph completion benchmarks. Jinfa Yang, Xianghua Ying, Yongjie Shi, Xin Tong 0007, Ruibin Wang, Taiyan Chen, Bowei Xing |
COLING | 3 |
| 2022 | Look Back and Forth: Video Super-Resolution with Explicit Temporal Difference ModelingabstractTemporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or complex motion, resulting in serious distortion and artifacts. In this paper, we propose to explore the role of explicit temporal difference modeling in both LR and HR space. Instead of directly feeding consecutive frames into a VSR model, we propose to compute the temporal difference between frames and divide those pixels into two subsets according to the level of difference. They are separately processed with two branches of different receptive fields in order to better extract complementary information. To further enhance the super-resolution result, not only spatial residual features are extracted, but the difference between consecutive frames in high-frequency domain is also computed. It allows the model to exploit intermediate SR results in both future and past to refine the current SR output. The difference at different time steps could be cached such that information from further distance in time could be propagated to the current frame for refinement. Experiments on several video super-resolution benchmark datasets demonstrate the effectiveness of the proposed method and its favorable performance against state-of-the-art methods. Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Ruihuang Li, Yongjie Shi, Huchuan Lu, Yu-Wing Tai |
CVPR | 6 |
| 2022 | Transformer Based Line Segment Classifier with Image Context for Real-Time Vanishing Point Detection in Manhattan WorldabstractPrevious works on vanishing point detection usually use geometric prior for line segment clustering. We find that image context can also contribute to accurate line classification. Based on this observation, we propose to classify line segments into three groups according to three unknown-but-sought vanishing points with Manhattan world assumption, using both geometric information and image context in this work. To achieve this goal, we propose a novel Transformer based Line segment Classifier (TLC) that can group line segments in images and estimate the corresponding vanishing points. In TLC, we design a line segment descriptor to represent line segments using their positions, directions and local image contexts. Transformer based feature fusion module is used to capture global features from all line segments, which is proved to improve the classification performance significantly in our experiments. By using a network to score line segments for outlier rejection, vanishing points can be got by Singular Value Decomposition (SVD) from the classified lines. The proposed method runs at 25 fps on one NVIDIA 2080Ti card for vanishing point detection. Experimental results on synthetic and real-world datasets demonstrate that our method is superior to other state-of-the-art methods on the balance between accuracy and efficiency, while keeping stronger generalization capability when trained and evaluated on different datasets. Xin Tong 0007, Xianghua Ying, Yongjie Shi, Ruibin Wang, Jinfa Yang |
CVPR | 3 |
| 2022 | Refactoring ISP for High-Level Vision TasksabstractThe image signal processing (ISP) pipeline, which transforms raw sensor measurement to a color image, is composed of a sequence of processing modules. Traditionally, the ISP pipeline is manually tuned by experts for human perception. The resulting handcrafted ISP configuration does not necessarily benefit the downstream high-level vision tasks. To mitigate these problems, this paper presents a simple yet effective framework based on Evolutionary Algorithm to search for a set of compact ISP configurations for high-level vision tasks. In particular, we encode ISP structure into a binary string and ISP parameters into a set of float numbers. Then we jointly optimize them with task-specific loss and ISP computation budgets (e.g., running time) through solving a nonlinear multi-objective optimization problem. By mutating the configurations of the ISP pipeline, we are able to remove redundant modules and design an ISP with both low cost and high accuracy. We validate the proposed method on extreme noisy and low-light raw images, and experimental results show that our framework can help find effective and efficient ISP configurations for both object detection and semantic segmentation tasks. We further provide a detailed analysis on the importance of different modules in the ISP configurations, which benefits the design of ISP for downstream tasks in the future. Yongjie Shi, Songjiang Li, Xu Jia 0012, Jianzhuang Liu |
ICRA | 1 |
| 2022 | Unsupervised Domain Adaptation for Semantic Segmentation of Urban Street Scenes Reflected by Convex MirrorsabstractReflective convex mirrors are often used on street corners or as passenger-side mirrors on cars to obtain scene information by reflecting blind spots in the field of view, which can provide safety for pedestrians and drivers on roads, driveways, and alleys that lack of visibility. In recent years, deep learning based scene understanding methods (e.g., semantic segmentation) have been rapidly developed. However, due to gaps in the geometric domain, models trained on normal images are not directly applicable to scenes with convex mirror reflections. In this paper, we propose a novel framework to reduce the domain gap between normal images and convex mirror reflection images. In particular, we geometrically model convex mirrors to obtain a differentiable convex mirror simulation layer, CMSL. With the help of CMSL, we perform adversarial domain adaptation on edges in the input space and semantic boundaries in the output space to reduce the geometric appearance gap between the synthetic and real images. To verify the effectiveness of our algorithm, we construct the first convex mirror reflection scene dataset CMR1K, which contains 268 images with fine annotations. Extensive experimental results show that our algorithm can significantly outperform the baseline and previous methods. For example, our method surpasses the baseline and AdvEnt by 10% and 3% in mIoU, respectively. Yongjie Shi, Xianghua Ying, Hongbin Zha |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Semi-Supervised Domain Adaptation Based on Dual-Level Domain Mixing for Semantic SegmentationabstractData-driven based approaches, in spite of great success in many tasks, have poor generalization when applied to unseen image domains, and require expensive cost of annotation especially for dense pixel prediction tasks such as semantic segmentation. Recently, both unsupervised domain adaptation (UDA) from large amounts of synthetic data and semi-supervised learning (SSL) with small set of labeled data have been studied to alleviate this issue. However, there is still a large gap on performance compared to their supervised counterparts. We focus on a more practical setting of semi-supervised domain adaptation (SSDA) where both a small set of labeled target data and large amounts of labeled source data are available. To address the task of SSDA, a novel framework based on dual-level domain mixing is proposed. The proposed framework consists of three stages. First, two kinds of data mixing methods are proposed to reduce domain gap in both region-level and sample-level respectively. We can obtain two complementary domain-mixed teachers based on dual-level mixed data from holistic and partial views respectively. Then, a student model is learned by distilling knowledge from these two teachers. Finally, pseudo labels of unlabeled data are generated in a self-training manner for another few rounds of teachers training. Extensive experimental results have demonstrated the effectiveness of our proposed framework on synthetic-to-real semantic segmentation benchmarks. Shuaijun Chen, Xu Jia 0012, Yongjie Shi, Jianzhuang Liu |
CVPR | 4 |
| 2021 | Multi-Target Domain Adaptation With Collaborative Consistency LearningabstractRecently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly extended to multiple target domains. In this work, we propose a collaborative learning framework to achieve unsupervised multi-target domain adaptation. An unsupervised domain adaptation expert model is first trained for each source-target pair and is further encouraged to collaborate with each other through a bridge built between different target domains. These expert models are further improved by adding the regularization of making the consistent pixel-wise prediction for each sample with the same structured context. To obtain a single model that works across multiple target domains, we propose to simultaneously learn a student model which is trained to not only imitate the output of each expert on the corresponding target domain, but also to pull different expert close to each other with regularization on their weights. Extensive experiments demonstrate that the proposed method can effectively exploit rich structured information contained in both labeled source domain and multiple unlabeled target domains. Not only does it perform well across multiple target domains but also performs favorably against state-of-the-art unsupervised domain adaptation methods specially trained on a single source-target pair. Code is available at https://github.com/junpan19/MTDA. Takashi Isobe, Xu Jia 0012, Shuaijun Chen, Yongjie Shi, Jianzhuang Liu, Huchuan Lu, Shengjin Wang |
CVPR | 5 |
| 2021 | Towards Cross-View Consistency in Semantic Segmentation While Varying View DirectionabstractSeveral images are taken for the same scene with many view directions. Given a pixel in any one image of them, its correspondences may appear in the other images. However, by using existing semantic segmentation methods, we find that the pixel and its correspondences do not always have the same inferred label as expected. Fortunately, from the knowledge of multiple view geometry, if we keep the position of a camera unchanged, and only vary its orientation, there is a homography transformation to describe the relationship of corresponding pixels in such images. Based on this fact, we propose to generate images which are the same as real images of the scene taken in certain novel view directions for training and evaluation. We also introduce gradient guided deformable convolution to alleviate the inconsistency, by learning dynamic proper receptive field from feature gradients. Furthermore, a novel consistency loss is presented to enforce feature consistency. Compared with previous approaches, the proposed method gets significant improvement in both cross-view consistency and semantic segmentation performance on images with abundant view directions, while keeping comparable or better performance on the existing datasets. Xin Tong 0007, Xianghua Ying, Yongjie Shi, He Zhao 0006, Ruibin Wang |
IJCAI | 3 |
| 2020 | RDCFace: Radial Distortion Correction for Face RecognitionabstractThe effects of radial lens distortion often appear in wide-angle cameras of surveillance and safeguard systems, which may severely degrade performances of previous face recognition algorithms. Traditional methods for radial lens distortion correction usually employ line features in scenarios that are not suitable for face images. In this paper, we propose a distortion-invariant face recognition system called RDCFace, which directly and only utilize the distorted images of faces, to alleviate the effects of radial lens distortion. RDCFace is an end-to-end trainable cascade network, which can learn rectification and alignment parameters to achieve a better face recognition performance without requiring supervision of facial landmarks and distortion parameters. We design sequential spatial transformer layers to optimize the correction, alignment, and recognition modules jointly. The feasibility of our method comes from implicitly using the statistics of the layout of face features learned from the large-scale face data. Extensive experiments indicate that our method is distortion robust and gains significant improvements on LFW, YTF, CFP, and RadialFace, a real distorted face benchmark compared with state-of-the-art methods. He Zhao 0006, Xianghua Ying, Yongjie Shi, Xin Tong 0007, Jingsi Wen, Hongbin Zha |
CVPR | 3 |
| 2020 | A Simple Yet Effective Pipeline For Radial Distortion CorrectionabstractEliminating the radial lens distortion of an image is a crucial preprocessing step for many computer vision applications. This paper explores a simple yet effective pipeline for radial distortion correction. Different from existing state-of-the-art methods that design complex network structure and concatenate multi-branch features. Our model uses a single network without any additional supervision. We design two differentiable layers to synthesize and rectify distorted images efficiently. Based on these layers, an online data synthesis strategy, a sampling grid loss, and an image reprojection loss are proposed to improve the distortion correction accuracy. Compared with the state-of-the-art methods, our model achieves the best rectification quality on both the synthetic and real distorted images with dozens of times faster inference speed. The training data and codes will be released.11https://github.com/MccreeZhao/RDCPipeline He Zhao 0006, Yongjie Shi, Xin Tong 0007, Xianghua Ying, Hongbin Zha |
ICIP | 2 |
| 2020 | Qamface: Quadratic Additive Angular Margin Loss For Face RecognitionabstractThe angular-based softmax losses and their variants achieve great success in face recognition based on deep learning. ArcFace [1] which directly maximize decision boundary in angular space is one of the most popular and effective loss function. In this paper, we analyze the inherent limitations of ArcFace, including the non-monotonic logit and gradient curve, and inappropriate trend of loss value. To address these problems, we propose a novel loss function named the Quadratic Additive Angular Margin Loss (QAMFace). It takes the value of the angle through a quadratic function rather than cosine function as the target logit. Our QAMFace is easy to implement and only adds negligible computational overhead. Experiments on several relevant benchmarks show that QAMFace performs better in convergence on feature embedding, and consistently outperforms the state-of-the-art face recognition methods. Our codes will be released soon.1 He Zhao 0006, Yongjie Shi, Xin Tong 0007, Xianghua Ying, Hongbin Zha |
ICIP | 2 |
| 2020 | Position-aware and Symmetry Enhanced GAN for Radial Distortion CorrectionabstractThis paper presents a novel method based on the generative adversarial network for radial distortion correction. Instead of generating a corrected image, our generator predicts a pixel flow map to measure the pixel offset between the distorted and corrected image. The quality of the generated pixel flow map and the warped image are judged by the discriminator. As texture far away from the image center has strong distortion, we develop an Adaptive Inverted Foveal layer which can transform the deformation to the intensity of the image to exploit this property. Rotation symmetry enhanced convolution kernels are applied to extract geometric features of different orientations explicitly. These learned features are recalibrated using the Squeeze-and-Excitation block to assign different weights for different directions. Moreover, we construct a first real-world radial distorted image dataset RD600 annotated with ground truth to evaluate our proposed method. We conduct extensive experiments to validate the effectiveness of each part of our framework. The further experiment shows our approach outperforms previous methods in both synthetic and real-world datasets quantitatively and qualitatively. Yongjie Shi, Xin Tong 0007, Jingsi Wen, He Zhao 0006, Xianghua Ying, Hongbin Zha |
ICPR | 1 |
| 2020 | G-FAN: Graph-Based Feature Aggregation Network for Video Face RecognitionabstractIn this paper, we propose a graph-based feature aggregation network (G-FAN) for video face recognition. Compared with the still image, video face recognition exhibits great challenges due to huge intra-class variability and high interclass ambiguity. To address this problem, our G-FAN first uses a Convolutional Neural Network to extract deep features for every input face of a subject. Then, we build an affinity graph based on the relationship between facial features and apply Graph Convolutional Network to generate fine-grained quality vectors for each frame. Finally, the features among multiple frames are adaptively aggregated into a discriminative vector to represent a video face. Different from previous works that take a single image as input, our G-FAN could utilize the correlation information between image pairs and aggregate a template of face images simultaneously. The experiments on video face recognition benchmarks, including YTF, IJB-A, and IJB-C show that: (i) G-FAN automatically learns to advocate high-quality frames while repelling low-quality ones. (ii) G-FAN significantly boosts recognition accuracy and outperforms other state-of-the-art aggregation methods. He Zhao 0006, Yongjie Shi, Xin Tong 0007, Jingsi Wen, Xianghua Ying, Hongbin Zha |
ICPR | 2 |
| 2020 | 3D Orientation Estimation and Vanishing Point Extraction from Single Panoramas Using Convolutional Neural Networkabstract3D orientation estimation is a key component of many important computer vision tasks such as autonomous navigation and 3D scene understanding. This paper presents a new CNN architecture to estimate the 3D orientation of an omnidirectional camera with respect to the world coordinate system from a single spherical panorama. To train the proposed architecture, we leverage a dataset of panoramas named VOP60K from Google Street View with labeled 3D orientation, including 50 thousand panoramas for training and 10 thousand panoramas for testing. Previous approaches usually estimate 3D orientation under pinhole cameras. However, for a panorama, due to its larger field of view, previous approaches cannot be suitable. In this paper, we propose an edge extractor layer to utilize the low-level and geometric information of panorama, an attention module to fuse different features generated by previous layers. A regression loss for two column vectors of the rotation matrix and classification loss for the position of vanishing points are added to optimize our network simultaneously. The proposed algorithm is validated on our benchmark, and experimental results clearly demonstrate that it outperforms previous methods. Yongjie Shi, Xin Tong 0007, Jingsi Wen, He Zhao 0006, Xianghua Ying, Hongbin Zha |
ICRA | 1 |
| 2019 | Three Orthogonal Vanishing Points Estimation in Structured Scenes Using Convolutional Neural NetworksabstractInferring 3D geometric cues is a crucial step, whereas vanishing point plays a very important role in image understanding from a single image of structured scenes. In this paper, we construct a 330-thousand-item image database of structured scenes labeled by vanishing points, focal length and camera orientation. We grab over 300 thousand Google Street View images which cover the downtown and neighboring areas of New York, Los Angeles, Chicago and etc. The prediction error is characterized by a loss function by imposing a regularization item derived from the geometric constraint of orthogonal vanishing points and focal length. Moreover, we collect about 30 thousand indoor images using a full 360-degree panorama camera taken by ourselves in room, office, library and etc. We also using Convolutional Neural Networks to transfer learning from street view images to indoor images. Extensive experiments demonstrate that our algorithm outperforms state-of-the-art non-learned approaches. Yongjie Shi, Danfeng Zhang, Jingsi Wen, Xin Tong 0007, He Zhao 0006, Xianghua Ying, Hongbin Zha |
ICIP | 1 |
| 2018 | Improve Cross-Domain Face Recognition with IBN-blockabstractDomain adaptation is one of the major challenges for face recognition (FR). Most large-scale FR training datasets are built from massive images crawled from the Internet, while in practical applications face images come from specific scenarios. Especially, for applications like ID card verification, registered face images are taken in controlled environments while probe face images are not. There are different distributions between source domain and target domain. In this paper, we propose to use Instance-Batch-Normalization (IBN) block to improve cross-domain FR performance. A million-scale cross-domain test set named IDCard-Scene-1M is used for evaluation. CNN models are trained with an improved loss function which we call L2-ASoftmax Loss. Without using any data from the target domain, IBN-block increased recall rates (@FPR = 10−6) on IDCard-Scene-1M by 1 percentage point for different CNN models. Besides, experiments show that the proposed IBN-CNN models trained with L2-ASoftmax Loss made state-of-the-art performance on MegaFace evaluation. Yangchun Qing, Yongjie Shi, Yining Lin |
IEEE BigData | 3 |
| 2018 | Radial Lens Distortion Correction by Adding a Weight Layer with Inverted Foveal Models to Convolutional Neural NetworksabstractRadial lens distortion often exists in images taken by commercial cameras, which does not satisfy the assumption of pinhole camera model. Eliminating the radial lens distortion of an image is necessary as a preprocessing step for many vision applications. Some paper has employed Convolutional Neural Networks (CNNs), to achieve radial distortion correction. They generated images with a large number of images of high variation of radial distortion, which can be well exploited by deep CNN with a high learning capacity, and reach the state-of-the-art results. In this paper, we claim that a weight layer with inverted foveal models can be added to these existing CNNs methods for radial distortion correction. In the widely used very deep Resnet-18 model, our method achieves about 20 percent decrease in the loss function with faster convergence compared to the previous methods. Yongjie Shi, Danfeng Zhang, Jingsi Wen, Xin Tong 0007, Xianghua Ying, Hongbin Zha |
ICPR | 1 |
| 2017 | Gourd pyrography art simulating based on non-photorealistic rendering
Wenhua Qian, Dan Xu 0001, Kun Yue, Yongjie Shi |
Multim. Tools Appl. | 6 |