Beiwen Tian

dblp:302/0648 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2024
0000-0002-2651-913XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 32% Segmentation and scene understanding · 23% Transfer learning and domain adaptation · 10%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth estimation
0.812024
Adaptive Surface Normal Constraint for Geometric Estimation From Monocular Images · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-target domain adaptation
0.812024
Training-Free Model Merging for Multi-target Domain Adaptation · ECCV (47) 2024
Computer vision › 3D vision
surface normal estimation
0.812024
Adaptive Surface Normal Constraint for Geometric Estimation From Monocular Images · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Efficient and distributed learning › model merging
training-free model merging
0.812024
Training-Free Model Merging for Multi-target Domain Adaptation · ECCV (47) 2024
Computer vision › 3D vision
3d object detection
0.712023
DQS3D: Densely-matched Quantization-aware Semi-supervised 3D Detection · ICCV 2023
Machine learning › Time series and sequential data › anomaly detection
anomaly segmentation
0.712023
Unsupervised Road Anomaly Detection with Language Anchors · ICRA 2023
Robotics › Autonomous driving
perception
0.712023
Unsupervised Road Anomaly Detection with Language Anchors · ICRA 2023
Computer vision › 3D vision › point cloud analysis
point cloud learning
0.712023
From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds · ICRA 2023
Robotics › Autonomous driving › autonomous driving perception
road anomaly detection
0.712023
Unsupervised Road Anomaly Detection with Language Anchors · ICRA 2023
Computer vision › 3D vision › 3d scene understanding
room layout estimation
0.712023
From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds · ICRA 2023
Computer vision › Segmentation and scene understanding
scene understanding
0.712023
From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds · ICRA 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
Delving into Shape-aware Zero-shot Semantic Segmentation · CVPR 2023
Computer vision › 3D vision › 3d object detection › label-efficient 3d object detection
semi-supervised 3d object detection
0.712023
DQS3D: Densely-matched Quantization-aware Semi-supervised 3D Detection · ICCV 2023
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
shape-based segmentation
0.712023
Delving into Shape-aware Zero-shot Semantic Segmentation · CVPR 2023
Machine learning › Time series and sequential data › anomaly detection
unsupervised anomaly detection
0.712023
Unsupervised Road Anomaly Detection with Language Anchors · ICRA 2023
Computer vision › Segmentation and scene understanding › semantic segmentation › open-vocabulary segmentation
zero-shot semantic segmentation
0.712023
Delving into Shape-aware Zero-shot Semantic Segmentation · CVPR 2023
Machine learning › Transfer learning and domain adaptation
cross-task generalization
0.612022
Unsupervised Cross-Task Generalization via Retrieval Augmentation · NeurIPS 2022
Computer vision › Segmentation and scene understanding
instance segmentation
0.612022
TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun Distillation · NeurIPS 2022
Machine learning › Learning paradigms › multi-task learning
multi-task language model
0.612022
Unsupervised Cross-Task Generalization via Retrieval Augmentation · NeurIPS 2022
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.612022
TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun Distillation · NeurIPS 2022
Information retrieval › retrieval models › neural retrieval
dense retrieval
0.612022
Unsupervised Cross-Task Generalization via Retrieval Augmentation · NeurIPS 2022
Information retrieval
retrieval augmentation
0.612022
Unsupervised Cross-Task Generalization via Retrieval Augmentation · NeurIPS 2022
Computer vision › 3D vision › spatial understanding
geometric context
0.212024
Adaptive Surface Normal Constraint for Geometric Estimation From Monocular Images · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Learning paradigms › semi-supervised learning
omni-supervised learning
0.212023
From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds · ICRA 2023
Computer vision › Vision and language
vision-language pretraining
0.212023
Delving into Shape-aware Zero-shot Semantic Segmentation · CVPR 2023

Methods — techniques the papers use, named apart from their topics

model merging · 0.8adaptive surface normal constraint · 0.8spectral methods · 0.7self-teaching · 0.7self-supervised features · 0.7quantization-aware training · 0.7quad set matching · 0.7pseudo-label harvesting · 0.7laplacian eigenvectors · 0.7exponential moving averaging · 0.7retrieval augmentation · 0.6pairwise reranking · 0.6
YearPublicationVenuePosition
2024 Training-Free Model Merging for Multi-target Domain Adaptation
Wenyi Li 0001, Huan-ang Gao, Mingju Gao, Beiwen Tian, Rong Zhi, Hao Zhao 0002
ECCV (47)4
2024 Adaptive Surface Normal Constraint for Geometric Estimation From Monocular Images
abstract
We introduce a novel approach to learn geometries such as depth and surface normal from images while incorporating geometric context. The difficulty of reliably capturing geometric context in existing methods impedes their ability to accurately enforce the consistency between the different geometric properties, thereby leading to a bottleneck of geometric estimation quality. We therefore propose the Adaptive Surface Normal (ASN) constraint, a simple yet efficient method. Our approach extracts geometric context that encodes the geometric variations present in the input image and correlates depth estimation with geometric constraints. By dynamically determining reliable local geometry from randomly sampled candidates, we establish a surface normal constraint, where the validity of these candidates is evaluated using the geometric context. Furthermore, our normal estimation leverages the geometric context to prioritize regions that exhibit significant geometric variations, which makes the predicted normals accurately capture intricate and detailed geometric information. Through the integration of geometric context, our method unifies depth and surface normal estimations within a cohesive framework, which enables the generation of high-quality 3D geometry from images. We validate the superiority of our approach over state-of-the-art methods through extensive evaluations and comparisons on diverse indoor and outdoor datasets, showcasing its efficiency and robustness.
Xiaoxiao Long, Yuhang Zheng 0004, Yupeng Zheng, Beiwen Tian, Cheng Lin 0001, Lingjie Liu, Hao Zhao 0002, Guyue Zhou, Wenping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Delving into Shape-aware Zero-shot Semantic Segmentation
abstract
Thanks to the impressive progress of large-scale vision-language pretraining, recent recognition models can classify arbitrary objects in a zero-shot and open-set manner, with a surprisingly high accuracy. However, translating this success to semantic segmentation is not trivial, because this dense prediction task requires not only accurate semantic understanding but also fine shape delineation and existing vision-language models are trained with image-level language descriptions. To bridge this gap, we pursue shape-aware zero-shot semantic segmentation in this study. Inspired by classical spectral methods in the image segmentation literature, we propose to leverage the eigen vectors of Laplacian matrices constructed with self-supervised pixel-wise features to promote shape-awareness. Despite that this simple and effective technique does not make use of the masks of seen classes at all, we demonstrate that it out-performs a state-of-the-art shape-aware formulation that aligns ground truth and predicted edges during training. We also delve into the performance gains achieved on different datasets using different backbones and draw several interesting and conclusive observations: the benefits of promoting shape-awareness highly relates to mask compactness and language embedding locality. Finally, our method sets new state-of-the-art performance for zero-shot semantic segmentation on both Pascal and COCO, with significant margins. Code and models will be accessed at SAZS.
Beiwen Tian, Kehua Sheng, Bo Zhang 0106, Hao Zhao 0002, Guyue Zhou
CVPR2
2023 DQS3D: Densely-matched Quantization-aware Semi-supervised 3D Detection
abstract
In this paper, we study the problem of semi-supervised 3D object detection, which is of great importance considering the high annotation cost for cluttered 3D indoor scenes. We resort to the robust and principled framework of self-teaching, which has triggered notable progress for semi-supervised learning recently. While this paradigm is natural for image-level or pixel-level prediction, adapting it to the detection problem is challenged by the issue of proposal matching. Prior methods are based upon two-stage pipelines, matching heuristically selected proposals generated in the first stage and resulting in spatially sparse training signals. In contrast, we propose the first semi-supervised 3D detection algorithm that works in the single-stage manner and allows spatially dense training signals. A fundamental issue of this new design is the quantization error caused by point-to-voxel discretization, which inevitably leads to misalignment between two transformed views in the voxel domain. To this end, we derive and implement closed-form rules that compensate this misalignment on-the-fly. Our results are significant, e.g., promoting Scan-Net [email protected] from 35.2% to 48.5% using 20% annotation. Codes and data are publicly available1.
Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou
ICCV2
2023 From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds
abstract
Room layout estimation is a long-existing robotic vision task that benefits both environment sensing and motion planning. However, layout estimation using point clouds (PCs) still suffers from data scarcity due to annotation difficulty. As such, we address the semi-supervised setting of this task based upon the idea of model exponential moving averaging. But adapting this scheme to the state-of-the-art (SOTA) solution for PC-based layout estimation is not straightforward. To this end, we define a quad set matching strategy and several consistency losses based upon metrics tailored for layout quads. Besides, we propose a new online pseudo-label harvesting algorithm that decomposes the distribution of a hybrid distance measure between quads and PC into two components. This technique does not need manual threshold selection and intuitively encourages quads to align with reliable layout points. Surprisingly, this framework also works for the fully-supervised setting, achieving a new SOTA on the ScanNet benchmark. Last but not least, we also push the semi-supervised setting to the realistic omni-supervised setting, demonstrating significantly promoted performance on a newly annotated ARKitScenes testing set. Our codes, data and models are made publicly available**Code: https://github.com/AIR-DISCOVER/Omni-PQ.
Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Yurong Chen 0001, Hongbin Zha
ICRA2
2023 Unsupervised Road Anomaly Detection with Language Anchors
abstract
Road anomaly detection is critical to safe autonomous driving, because current road scene understanding models are usually trained in a closed-set manner and fail to identify unknown objects. What's worse, it is difficult, if not impossible, to collect a large-scale dataset with anomaly annotations. So this paper studies unsupervised anomaly detection which finds out anomaly regions using scene parsing logits solely. While former methods depend on the weights learned from the closed training set as anchors for logit generation, we resort to language anchors that are learned from enormous paired vision and language data. Thanks to rich open-set semantic information contained in these language anchors, our method performs better than former unsupervised counterparts while maintaining the advantage of training without accessing any out-of-distribution data. We delve into this new paradigm and identify the superiority of using pair-wise binary logits, which we credit to a better understanding of the negation language anchor. Last but not least, we find that the former top-1 selection of semantic labels for uncertainty measurement is problematic in many cases and a new blended standardization strategy brings clear improvements to our solution. We report state-of-the-art performance on FS LostAndFound, LostAndFound and RoadAnomaly datasets among comparable methods. The codes are publicly available at https://github.com/TB5z035/URAD-LA.git
Beiwen Tian, Mingdao Liu, Huan-ang Gao, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou
ICRA1
2022 TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun Distillation
abstract
Current referring expression comprehension algorithms can effectively detect or segment objects indicated by nouns, but how to understand verb reference is still under-explored. As such, we study the challenging problem of task oriented detection, which aims to find objects that best afford an action indicated by verbs like sit comfortably on. Towards a finer localization that better serves downstream applications like robot interaction, we extend the problem into task oriented instance segmentation. A unique requirement of this task is to select preferred candidates among possible alternatives. Thus we resort to the transformer architecture which naturally models pair-wise query relationships with attention, leading to the TOIST method. In order to leverage pre-trained noun referring expression comprehension models and the fact that we can access privileged noun ground truth during training, a novel noun-pronoun distillation framework is proposed. Noun prototypes are generated in an unsupervised manner and contextual pronoun features are trained to select prototypes. As such, the network remains noun-agnostic during inference. We evaluate TOIST on the large-scale task oriented dataset COCO-Tasks and achieve +10.7% higher $\rm{mAP^{box}}$ than the best-reported results. The proposed noun-pronoun distillation can boost $\rm{mAP^{box}}$ and $\rm{mAP^{mask}}$ by +2.6% and +3.6%. Codes and models are publicly available.
Pengfei Li 0007, Beiwen Tian, Yongliang Shi, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang
NeurIPS2
2022 Unsupervised Cross-Task Generalization via Retrieval Augmentation
abstract
Humans can perform unseen tasks by recalling relevant skills acquired previously and then generalizing them to the target tasks, even if there is no supervision at all. In this paper, we aim to improve this kind of cross-task generalization ability of massive multi-task language models, such as T0 and FLAN, in an unsupervised setting. We propose a retrieval-augmentation method named ReCross that takes a few unlabelled examples as queries to retrieve a small subset of upstream data and uses them to update the multi-task model for better generalization. ReCross is a straightforward yet effective retrieval method that combines both efficient dense retrieval and effective pair-wise reranking. Our results and analysis show that it significantly outperforms both non-retrieval methods and other baseline methods.
Bill Y. Lin, Kangmin Tan, Beiwen Tian, Xiang Ren 0001
NeurIPS4