EDBT 2026 Demo / reviewers in the wild / expert
Ruibin Wang
dblp:278/3974
· DBLP profile ↗
19ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Capability of large language models in assisting GPs with diagnoses
Ruibin Wang, Abdul Rehman 0007, Rupert Page, Hailing Li, Xiaokun Wang 0001, Xiaosong Yang, Jian J. Zhang 0001 |
Appl. Intell. | 1 |
| 2026 | MVFormer: Multi-View Point Cloud Transformer for 3D Mechanical Component Recognition
Ruibin Wang, Xianghua Ying, Bowei Xing |
Int. J. Comput. Vis. | 1 |
| 2025 | Improving Bird's Eye View based 3D object detection via Learnable Perspective View
Ruibin Wang, Bowei Xing, Xianghua Ying |
Expert Syst. Appl. | 1 |
| 2025 | Multi-modal Prompt Alignment with Fine-grained LLM Knowledge for Unsupervised Domain Adaptation
Bowei Xing, Xianghua Ying, Ruibin Wang, Ruohao Guo |
Int. J. Comput. Vis. | 3 |
| 2024 | VPDETR: End-to-End Vanishing Point DEtection TRansformersabstractIn the field of vanishing point detection, previous works commonly relied on extracting and clustering straight lines or classifying candidate points as vanishing points. This paper proposes a novel end-to-end framework, called VPDETR (Vanishing Point DEtection TRansformer), that views vanishing point detection as a set prediction problem, applicable to both Manhattan and non-Manhattan world datasets. By using the positional embedding of anchor points as queries in Transformer decoders and dynamically updating them layer by layer, our method is able to directly input images and output their vanishing points without the need for explicit straight line extraction and candidate points sampling. Additionally, we introduce an orthogonal loss and a cross-prediction loss to improve accuracy on the Manhattan world datasets. Experimental results demonstrate that VPDETR achieves competitive performance compared to state-of-the-art methods, without requiring post-processing. Taiyan Chen, Xianghua Ying, Jinfa Yang, Ruibin Wang, Ruohao Guo, Bowei Xing, Ji Shi 0003 |
AAAI | 4 |
| 2024 | Hierarchical Unsupervised Relation Distillation for Source Free Domain Adaptation
Bowei Xing, Xianghua Ying, Ruibin Wang, Ruohao Guo, Ji Shi 0003, Wenzhen Yue |
ECCV (50) | 3 |
| 2024 | Masked Local-Global Representation Learning for 3D Point Cloud Domain AdaptationabstractPoint cloud is a popular and widely used geometric representation, which has attracted significant attention in 3D vision. However, the geometric variability of point cloud representations across different datasets can cause domain discrepancies, which hinder knowledge transfer and model generalization, resulting in degraded performance in target domain. In this paper, we present a novel approach to improve point cloud domain adaptation by employing masked representation learning in a self-supervised manner. Specifically, our method combines masked feature prediction and masked sample consistency to encode both local structure and global semantic information for learning invariant point cloud representation across domains. Moreover, to learn domain-specific representation and transfer knowledge from source to target, we propose prototype-calibrated self-training. By exploiting class-wise prototypes in the shared feature space, the soft pseudo labels can be adaptively denoised, which benefits the decision boundary learning in target domain. We conduct experiments on PointDA-10 and PointSegDA for 3D point cloud shape classification and semantic segmentation, respectively. The results demonstrate the effectiveness of our method and show that we can achieve the new state-of-the-art performance on point cloud domain adaptation. Bowei Xing, Xianghua Ying, Ruibin Wang |
ICRA | 3 |
| 2024 | Rectifying self-training with neighborhood consistency and proximity for source-free domain adaptation
Bowei Xing, Xianghua Ying, Ruibin Wang |
Neurocomputing | 3 |
| 2024 | Exploiting Temporal Correlations for 3D Human Pose EstimationabstractExploiting the rich temporal information in human pose sequences to facilitate 3D pose estimation has garnered particular attention. While various learning architectures have been designed for temporal exploiting, these architectures are usually trained via the 3D pose loss independently imposed on every single frame, without explicit temporal signals introduced for supervision. This inevitably increases the difficulty of temporal exploiting, since the network must reason about the meaningful temporal information based on the non-temporal single-frame supervision first. Only then, the network can utilize this information to guide sequence modeling. Recently, some work introduce temporal smoothness as an explicit supervision signal, which makes the network more straightforwardly reaches the temporal information from the supervision signal, thus improving the temporal exploiting. However, the temporal smoothness only roughly measures the short-term temporal properties between adjacent frame pairs. In this work, we propose to generalize the supervision of temporal smoothness to temporal correlations, letting the network precisely consider more comprehensive temporal properties in sequences. We contribute two novel correlation-based loss functions, which adopt different strategies to respectively regularize the encoder and decoder sides of the network for temporal exploiting. Besides, we design a pre-training scheme to ensure a general convergence of existing pose estimators under our correlation losses. Experiments on three benchmarks demonstrate that our method can be compatible with different networks, improving their temporal exploiting ability to output more accurate and robust pose estimations. Ruibin Wang, Xianghua Ying, Bowei Xing |
IEEE Trans. Multim. | 1 |
| 2024 | Improving static and temporal knowledge graph embedding using affine transformations of entitiesabstractTo find a suitable embedding for a knowledge graph (KG) remains a big challenge nowadays. By measuring the distance or plausibility of triples and quadruples in static and temporal knowledge graphs, many reliable knowledge graph embedding (KGE) models are proposed. However, these classical models may not be able to represent and infer various relation patterns well, such as TransE cannot represent symmetric relations, DistMult cannot represent inverse relations, RotatE cannot represent multiple relations, etc.. In this paper, we improve the ability of these models to represent various relation patterns by introducing the affine transformation framework. Specifically, we first utilize a set of affine transformations related to each relation or timestamp to operate on entity vectors, and then these transformed vectors can be applied not only to static KGE models, but also to temporal KGE models. The main advantage of using affine transformations is their good geometry properties with interpretability. Our experimental results demonstrate that the proposed intuitive design with affine transformations provides a statistically significant increase in performance with adding a few extra processing steps and keeping the same number of embedding parameters. Taking TransE as an example, we employ the scale transformation (the special case of an affine transformation). Surprisingly, it even outperforms RotatE to some extent on various datasets. We also introduce affine transformations into RotatE, Distmult, ComplEx, TTransE and TComplEx respectively, and experiments demonstrate that affine transformations consistently and significantly improve the performance of state-of-the-art KGE models on both static and temporal knowledge graph benchmarks. Jinfa Yang, Xianghua Ying, Yongjie Shi, Ruibin Wang |
J. Web Semant. | 4 |
| 2023 | ECO-3D: Equivariant Contrastive Learning for Pre-training on Perturbed 3D Point CloudabstractIn this work, we investigate contrastive learning on perturbed point clouds and find that the contrasting process may widen the domain gap caused by random perturbations, making the pre-trained network fail to generalize on testing data. To this end, we propose the Equivariant COntrastive framework which closes the domain gap before contrasting, further introduces the equivariance property, and enables pre-training networks under more perturbation types to obtain meaningful features. Specifically, to close the domain gap, a pre-trained VAE is adopted to convert perturbed point clouds into less perturbed point embedding of similar domains and separated perturbation embedding. The contrastive pairs can then be generated by mixing the point embedding with different perturbation embedding. Moreover, to pursue the equivariance property, a Vector Quantizer is adopted during VAE training, discretizing the perturbation embedding into one-hot tokens which indicate the perturbation labels. By correctly predicting the perturbation labels from the perturbed point cloud, the property of equivariance can be encouraged in the learned features. Experiments on synthesized and real-world perturbed datasets show that ECO-3D outperforms most existing pre-training strategies under various downstream tasks, achieving SOTA performance for lots of perturbations. Ruibin Wang, Xianghua Ying, Bowei Xing, Jinfa Yang |
AAAI | 1 |
| 2023 | Cross-Modal Contrastive Learning for Domain Adaptation in 3D Semantic SegmentationabstractDomain adaptation for 3D point cloud has attracted a lot of interest since it can avoid the time-consuming labeling process of 3D data to some extent. A recent work named xMUDA leveraged multi-modal data to domain adaptation task of 3D semantic segmentation by mimicking the predictions between 2D and 3D modalities, and outperformed the previous single modality methods only using point clouds. Based on it, in this paper, we propose a novel cross-modal contrastive learning scheme to further improve the adaptation effects. By employing constraints from the correspondences between 2D pixel features and 3D point features, our method not only facilitates interaction between the two different modalities, but also boosts feature representations in both labeled source domain and unlabeled target domain. Meanwhile, to sufficiently utilize 2D context information for domain adaptation through cross-modal learning, we introduce a neighborhood feature aggregation module to enhance pixel features. The module employs neighborhood attention to aggregate nearby pixels in the 2D image, which relieves the mismatching between the two different modalities, arising from projecting relative sparse point cloud to dense image pixels. We evaluate our method on three unsupervised domain adaptation scenarios, including country-to-country, day-to-night, and dataset-to-dataset. Experimental results show that our approach outperforms existing methods, which demonstrates the effectiveness of the proposed method. Bowei Xing, Xianghua Ying, Ruibin Wang, Jinfa Yang, Taiyan Chen |
AAAI | 3 |
| 2023 | A Multi-Frame Rate Network with Attention Mechanism for Depression Severity EstimationabstractThe diagnosis of depression mainly depends on clinicians’ experience and questionnaire results, making it a time-consuming and subjective process that demands significant allocation of human resources. Numerous automatic depression estimation (ADE) systems, based on facial cues, have been introduced to estimate the severity and assist clinicians in diagnosis. However, traditional methods adopt a single sampling frame rate, which makes it leads to a tradeoff between the loss of critical vision information and calculation redundancy. In this paper, we propose a Multi-Frame Rate Attention Convolutional Neural Network (MFRA) to effectively mine the facial cues of patients for estimating the severity of depression. Specifically, we adpot a two-branch network structure, in which one branch uses high frame sampling rate to capture more nuances of facial changes, and the other uses low frame sampling rate to focus more on the spatial information of the video. Furthermore, considering the manually selected local facial features will introduce noise, we introduce attention modules to make MFRA concentrate on the facial regions related to depression. Finally, the feature vectors extracted from the two branches are aggregated to output the severity of depression. The experimental results on two datasets, AVEC 2013 and AVEC 2014, show that this method can effectively capture spatiotemporal features, and the prediction results are superior to most video-based depression prediction methods. Ruibin Wang, Jiashun Wang, Yun Yang 0003 |
BIBM | 1 |
| 2023 | Improving point cloud classification and segmentation via parametric veronese mappingabstractDeep learning based 3D point cloud classification and segmentation has achieved remarkable success. Existing methods are usually implemented in the original space with 3D coordinates as inputs. However, we find that point networks taking only information of first-order coordinates hardly learn geometric features of higher order, such as point cloud normals or poses. In this study, we propose to map the input point clouds into a non-linear space to facilitate networks learning and leveraging high-order features. Firstly, we design the Parametric Veronese Mapping (PVM) function which automatically learns to map point clouds into a non-linear space. As a result, the mapped point clouds are enriched with high-order elements and maintain the basic point set properties as in the original 3D space. We can then exploit existing networks to learn high-order features from mapped point clouds. Secondly, we contribute a two-stage transformation learning module that modifies the previous one-stage module to better leverage high-order features for aligning point clouds in the projective space. Finally, an interaction module is designed to learn more discriminative features by aggregating information from both the original and projective space. Extensive experiments demonstrate that our method successfully improves the ability of most existing networks to learn high-order features and thus contributing to more accurate classification and segmentation. Moreover, the resulting models show stronger robustness to affine transformations and real-world perturbations. Ruibin Wang, Xianghua Ying, Bowei Xing, Xin Tong 0007, Taiyan Chen, Jinfa Yang, Yongjie Shi |
Pattern Recognit. | 1 |
| 2022 | Learning Hierarchy-Aware Quaternion Knowledge Graph Embeddings with Representing Relations as 3D RotationsabstractKnowledge graph embedding aims to represent entities and relations as low-dimensional vectors, which is an effective way for predicting missing links. It is crucial for knowledge graph embedding models to model and infer various relation patterns, such as symmetry/antisymmetry. However, many existing approaches fail to model semantic hierarchies, which are common in the real world. We propose a new model called HRQE, which represents entities as pure quaternions. The relational embedding consists of two parts: (a) Using unit quaternions to represent the rotation part in 3D space, where the head entities are rotated by the corresponding relations through Hamilton product. (b) Using scale parameters to constrain the modulus of entities to make them have hierarchical distributions. To the best of our knowledge, HRQE is the first model that can encode symmetry/antisymmetry, inversion, composition, multiple relation patterns and learn semantic hierarchies simultaneously. Experimental results demonstrate the effectiveness of HRQE against some of the SOTA methods on four well-established knowledge graph completion benchmarks. Jinfa Yang, Xianghua Ying, Yongjie Shi, Xin Tong 0007, Ruibin Wang, Taiyan Chen, Bowei Xing |
COLING | 5 |
| 2022 | Transformer Based Line Segment Classifier with Image Context for Real-Time Vanishing Point Detection in Manhattan WorldabstractPrevious works on vanishing point detection usually use geometric prior for line segment clustering. We find that image context can also contribute to accurate line classification. Based on this observation, we propose to classify line segments into three groups according to three unknown-but-sought vanishing points with Manhattan world assumption, using both geometric information and image context in this work. To achieve this goal, we propose a novel Transformer based Line segment Classifier (TLC) that can group line segments in images and estimate the corresponding vanishing points. In TLC, we design a line segment descriptor to represent line segments using their positions, directions and local image contexts. Transformer based feature fusion module is used to capture global features from all line segments, which is proved to improve the classification performance significantly in our experiments. By using a network to score line segments for outlier rejection, vanishing points can be got by Singular Value Decomposition (SVD) from the classified lines. The proposed method runs at 25 fps on one NVIDIA 2080Ti card for vanishing point detection. Experimental results on synthetic and real-world datasets demonstrate that our method is superior to other state-of-the-art methods on the balance between accuracy and efficiency, while keeping stronger generalization capability when trained and evaluated on different datasets. Xin Tong 0007, Xianghua Ying, Yongjie Shi, Ruibin Wang, Jinfa Yang |
CVPR | 4 |
| 2022 | ART-Point: Improving Rotation Robustness of Point Cloud Classifiers via Adversarial RotationabstractPoint cloud classifiers with rotation robustness have been widely discussed in the 3D deep learning community. Most proposed methods either use rotation invariant descriptors as inputs or try to design rotation equivariant networks. However, robust models generated by these methods have limited performance under clean aligned datasets due to modifications on the original classifiers or input space. In this study, for the first time, we show that the rotation robustness of point cloud classifiers can also be acquired via adversarial training with better performance on both rotated and clean datasets. Specifically, our proposed framework named ART-Point regards the rotation of the point cloud as an attack and improves rotation robustness by training the classifier on inputs with Adversarial RoTations. We contribute an axis-wise rotation attack that uses back-propagated gradients of the pre-trained model to effectively find the adversarial rotations. To avoid model overfitting on adversarial inputs, we construct rotation pools that leverage the transferability of adversarial rotations among samples to increase the diversity of training data. Moreover, we propose a fast one-step optimization to efficiently reach the final robust model. Experiments show that our proposed rotation attack achieves a high success rate and ART-Point can be used on most existing classifiers to improve the rotation robustness while showing better performance on clean datasets than state-of-the-art methods. Ruibin Wang, Dacheng Tao |
CVPR | 1 |
| 2021 | Towards Cross-View Consistency in Semantic Segmentation While Varying View DirectionabstractSeveral images are taken for the same scene with many view directions. Given a pixel in any one image of them, its correspondences may appear in the other images. However, by using existing semantic segmentation methods, we find that the pixel and its correspondences do not always have the same inferred label as expected. Fortunately, from the knowledge of multiple view geometry, if we keep the position of a camera unchanged, and only vary its orientation, there is a homography transformation to describe the relationship of corresponding pixels in such images. Based on this fact, we propose to generate images which are the same as real images of the scene taken in certain novel view directions for training and evaluation. We also introduce gradient guided deformable convolution to alleviate the inconsistency, by learning dynamic proper receptive field from feature gradients. Furthermore, a novel consistency loss is presented to enforce feature consistency. Compared with previous approaches, the proposed method gets significant improvement in both cross-view consistency and semantic segmentation performance on images with abundant view directions, while keeping comparable or better performance on the existing datasets. Xin Tong 0007, Xianghua Ying, Yongjie Shi, He Zhao 0006, Ruibin Wang |
IJCAI | 5 |
| 2020 | Sketch-based modeling with a differentiable rendererabstractAbstract Sketch‐based modeling aims to recover three‐dimensional (3D) shape from two‐dimensional line drawings. However, due to the sparsity and ambiguity of the sketch, it is extremely challenging for computers to interpret line drawings of physical objects. Most conventional systems are restricted to specific scenarios such as recovering for specific shapes, which are not conducive to generalize. Recent progress of deep learning methods have sparked new ideas for solving computer vision and pattern recognition issues. In this work, we present an end‐to‐end learning framework to predict 3D shape from line drawings. Our approach is based on a two‐steps strategy, it converts the sketch image to its normal image, then recover the 3D shape subsequently. A differentiable renderer is proposed and incorporated into this framework, it allows the integration of the rendering pipeline with neural networks. Experimental results show our method outperforms the state‐of‐art, which demonstrates that our framework is able to cope with the challenges in single sketch‐based 3D shape modeling. Ruibin Wang, Tao Jiang 0020, Li Wang 0105, Yanran Li, Xiaosong Yang, Jian J. Zhang 0001 |
Comput. Animat. Virtual Worlds | 2 |