VLDB 2026 Research / reviewers in the wild / expert
Haotian Dong
dblp:314/2870
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 56% Generative modeling · 20% Autonomous driving · 13% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 75% Image and video processing · 25% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 77% Cloud and datacenter computing · 23% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration · ICCV 2025 |
Machine learning › Generative modeling › synthetic data generation
LLM-based data generation |
0.9 | 1 | 2025 | Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous Driving · ICLR 2025 |
Computer vision › 3D vision
point cloud analysis |
0.9 | 1 | 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completion · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › point cloud processing
point cloud completion |
0.9 | 1 | 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completion · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud convolution |
0.9 | 1 | 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completion · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › 3d shape reconstruction
shape completion |
0.9 | 1 | 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completion · IEEE Trans. Multim. 2025 |
Robotics › Autonomous driving
trajectory prediction |
0.9 | 1 | 2025 | Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous Driving · ICLR 2025 |
Visual content generation and editing
image and video generation |
0.9 | 1 | 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration · ICCV 2025 |
Image and video processing › image restoration
image inpainting |
0.9 | 1 | 2025 | HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image Inpainting · ACM Trans. Graph. 2025 |
Visual content generation and editing › video generation
multi-view video generation |
0.9 | 1 | 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration · ICCV 2025 |
Visual content generation and editing
video generation |
0.9 | 1 | 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration · ICCV 2025 |
Distributed systems › distributed machine learning
distributed training |
0.9 | 1 | 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM Training · EMNLP 2025 |
Computer vision › 3D vision
3d scene understanding |
0.7 | 1 | 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion · ICCV 2023 |
Machine learning › Deep learning architectures and training › transformer
multi-view transformer |
0.7 | 1 | 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion · ICCV 2023 |
Computer vision › 3D vision › 3d scene understanding
semantic scene completion |
0.7 | 1 | 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion · ICCV 2023 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM Training · EMNLP 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.3 | 1 | 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM Training · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
noise decomposition · 1.7noise collaboration · 1.7synthetic data generation · 0.9sensor-aware data augmentation · 0.9point-block carving · 0.9large language model · 0.9joint distribution modeling · 0.9generative sub-networks · 0.9adaptive fusion kernels · 0.9multi-view feature synthesis · 0.7cross-view transformer · 0.73d convolutional kernel rotation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingabstractThe emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations.Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities.As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches.We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements.We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training. Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Zhi Wang 0001 |
EMNLP | 1 |
| 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration
Haotian Dong, Xin Wang 0118, Di Lin 0002, Yipeng Wu, Kairui Yang, Ping Li 0016, Qing Guo 0005 |
ICCV | 1 |
| 2025 | Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous DrivingabstractVehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using Large Language Models (LLMs) to generate abundant and realistic trajectories of interacting vehicles efficiently. These models rely on textual descriptions of vehicle-to-vehicle interactions on a map to produce the trajectories. We introduce Trajectory-LLM (Traj-LLM), a new approach that takes brief descriptions of vehicular interactions as input and generates corresponding trajectories. Unlike language-based approaches that translate text directly to trajectories, Traj-LLM uses reasonable driving behaviors to align the vehicle trajectories with the text. This results in an "interaction-behavior-trajectory" translation process. We have also created a new dataset, Language-to-Trajectory (L2T), which includes 240K textual descriptions of vehicle interactions and behaviors, each paired with corresponding map topologies and vehicle trajectory segments. By leveraging the L2T dataset, Traj-LLM can adapt interactive trajectories to diverse map topologies. Furthermore, Traj-LLM generates additional data that enhances downstream prediction models, leading to consistent performance improvements across public benchmarks. The source code is released at https://github.com/TJU-IDVLab/Traj-LLM. Kairui Yang, Gengjie Lin, Haotian Dong, Yipeng Wu, Die Zuo, Jibin Peng, Ziyuan Zhong, Xin Wang 0118, Qing Guo 0005, Xiaosong Jia, Junchi Yan, Di Lin 0002 |
ICLR | 4 |
| 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completionabstract3D point cloud completion is very challenging because it relies on accurately understanding the complex 3D shapes (e.g., high-curvature, concave/convex, and hollowed-out 3D shapes) and the unknown & diverse patterns of the partially available point clouds. In this paper, we propose a novel solution, i.e.,Point-block Carving(PC), for completing the complex 3D point cloud completion. Given the partial point cloud as the guidance, we carve a 3D block that contains the uniformly distributed 3D points, yielding the entire point cloud. We propose a new network architecture to achieve PC, i.e.,CarveNet. This network conducts the exclusive convolution on each block point, where the convolutional kernels are trained on the 3D shape data. CarveNet determines which point should be carved to recover the complete shapes' details effectively. Furthermore, we propose a sensor-aware method for data augmentation, i.e.,SensorAug, for training CarveNet on richer patterns of partial point clouds, thus enhancing the completion power of the network. The extensive evaluations on the ShapeNet, ShapNet-55/34 and KITTI datasets demonstrate the generality of our approach on the partial point clouds with diverse patterns. On these datasets, CarveNet successfully outperforms the state-of-the-art methods. Qing Guo 0005, Zhijie Wang 0014, Lubo Wang, Haotian Dong, Felix Juefei-Xu, Di Lin 0002, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image InpaintingabstractMulti-domain image inpainting utilizes complementary contextual information from auxiliary domain images to restore corrupted regions. While existing methods reconstruct auxiliary images to provide additional guidance, they face fundamental limitations: recovered pixels with complex patterns often lack representative details, while oversimplified patterns offer insufficient contextual information. To address these challenges, we propose HRC-Net, a novel framework incorporating three generative sub-networks for the comprehensive image inpainting task. Our architecture consists of: (1) A Hypothesis Sub-network that enables robust samplings of pixel-wise hypotheses from multi-domain inputs; (2) A Representative Sub-network that learns to score hypothesis quality based on contextual relevance; and (3) a Collaboration Sub-network that optimizes adaptive fusion kernels to integrate the most pertinent details. Together, these components model the joint distribution of representative scores and convolutional kernels, fostering a precise interaction between auxiliary hypotheses and target image corruption to meticulously repair the target image. Extensive evaluations across multiple benchmark datasets demonstrate HRC-Net's superior performance, significantly outperforming state-of-the-art methods in both quantitative metrics and visual quality. Xin Wang 0118, Di Lin 0002, Wanchao Su, Ji Du, Jie Zhang 0090, Haotian Dong, Ke Xu 0010, Qing Guo 0005, Ping Li 0016 |
ACM Trans. Graph. | 7 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 1 |
| 2023 | Expectile regression forest: A new nonparametric expectile regression modelabstractAbstract Classical nonlinear expectile regression has two shortcomings. It is difficult to choose a nonlinear function, and it does not consider the interaction effects among explanatory variables. Therefore, we combine the random forest model with the expectile regression method to propose a new nonparametric expectile regression model: expectile regression forest (ERF). The major novelty of the ERF model is using the bagging method to build multiple decision trees, calculating the conditional expectile of each leaf node in each decision tree, and deriving final results through aggregating these decision tree results via simple average approach. At the same time, in order to compensate for the black box problem in the model interpretation of the ERF model, the measurement of the importance of explanatory variable and the partial dependence is defined to evaluate the magnitude and direction of the influence of each explanatory variable on the response variable. The advantage of ERF model is illustrated by Monte Carlo simulation studies. The numerical simulation results show that the estimation and prediction ability of the ERF model is significantly better than alternative approaches. We also apply the ERF model to analyse the real data. From the nonparametric expectile regression analysis of these data sets, we have several conclusions that are consistent with the results of numerical simulation. Haotian Dong |
Expert Syst. J. Knowl. Eng. | 2 |