Haixin Shi

dblp:331/5873 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
object pose estimation
0.912025
Free-Moving Object Reconstruction and Pose Estimation with Virtual Camera · AAAI 2025
Computer vision › 3D vision › 3d reconstruction
object reconstruction
0.912025
Free-Moving Object Reconstruction and Pose Estimation with Virtual Camera · AAAI 2025
Computer vision › 3D vision
pose estimation
0.912025
Free-Moving Object Reconstruction and Pose Estimation with Virtual Camera · AAAI 2025
Computer vision › 3D vision
implicit neural representation
0.312025
Free-Moving Object Reconstruction and Pose Estimation with Virtual Camera · AAAI 2025

Methods — techniques the papers use, named apart from their topics

virtual camera · 0.9implicit neural representation · 0.9global optimization · 0.9
YearPublicationVenuePosition
2025 Free-Moving Object Reconstruction and Pose Estimation with Virtual Camera
abstract
We propose an approach for reconstructing free-moving object from a monocular RGB video. Most existing methods either assume scene prior, hand pose prior, object category pose prior, or rely on local optimization with multiple sequence segments. We propose a method that allows free interaction with the object in front of a moving camera without relying on any prior, and optimizes the sequence globally without any segments. We progressively optimize the object shape and pose simultaneously based on an implicit neural representation. A key aspect of our method is a virtual camera system that reduces the search space of the optimization significantly. We evaluate our method on the standard HO3D dataset and a collection of egocentric RGB sequences captured with a head-mounted device. We demonstrate that our approach outperforms most methods significantly, and is on par with recent techniques that assume prior information.
Haixin Shi, Yinlin Hu, Daniel Koguciuk, Juan-Ting Lin, Mathieu Salzmann, David Ferstl
AAAI1
2024 Adaptive Point Cloud Clustering Algorithm for Practical Roadside MmWave Radar Systems
abstract
Millimeter-Wave radar has been widely applied in the field of autonomous driving due to an excellent performance under complex weather conditions. However, in practical roadside scenarios, the challenge of sparse point clouds leading to clustering difficulties and the issue of large vehicle point clouds dispersing, resulting in fragmentation, currently hampers the practical ap-plication of radar sensors. We propose an adaptive point cloud clustering algorithm based on DBSCAN. First, we propose an improved DBSCAN clustering algorithm based on distance and speed thresholds, which enhances the differentiation of point clouds between different vehicles, and an adaptive ellipse gate strategy to solve the large vehicle point clouds fragmentation problem. Then, a secondary clustering algorithm based on azimuth is exploited, effectively addressing the issues of large vehicle fragmentation and anomalous speed values. Practical roadside experimental results demonstrate that our proposed algorithm significantly outperforms traditional algorithms, showing considerable potential in practical applications.
Luyi Zhang, Jinhang Zhang, Haixin Shi, Rui Chen 0001
VTC Spring3
2023 Two-level Data Augmentation for Calibrated Multi-view Detection
abstract
Data augmentation has proven its usefulness to improve model generalization and performance. While it is commonly applied in computer vision application when it comes to multi-view systems, it is rarely used. Indeed geometric data augmentation can break the alignment among views. This is problematic since multi-view data tend to be scarce and it is expensive to annotate.In this work we propose to solve this issue by introducing a new multi-view data augmentation pipeline that preserves alignment among views. Additionally to traditional augmentation of the input image we also propose a second level of augmentation applied directly at the scene level. When combined with our simple multi-view detection model, our two-level augmentation pipeline outperforms all existing baselines by a significant margin on the two main multi-view multi-person detection datasets WILD-TRACK and MultiviewX.
Martin Engilberge, Haixin Shi, Zhiye Wang, Pascal Fua
WACV2