EDBT 2026 Demo / reviewers in the wild / expert
Tianhao Fu
dblp:300/9624
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two Heads Are Better than One: Distilling Large Language Model Features into Small Models with Feature Decomposition and MixtureabstractMarket making (MM) through Reinforcement Learning (RL) has attracted significant attention in financial trading. With the development of Large Language Models (LLMs), more and more attempts are being made to apply LLMs to financial areas. A simple, direct application of LLM as an agent shows significant performance. Such methods are hindered by their slow inference speed, while most of the current research has not studied LLM distillation for this specific task. To address this, we first propose the normalized fluorescent probe to study the mechanism of the LLM’s feature. Based on the observation found by our investigation, we propose Cooperative Market Making (CMM), a novel framework that decouples LLM features across three orthogonal dimensions: layer, task, and data. Various student models collaboratively learn simple LLM features along with different dimensions, with each model responsible for a distinct feature to achieve knowledge distillation. Furthermore, CMM introduces an Hájek-MoE to integrate the output of the student models by investigating the contribution of different models in a kernel function-generated common feature space. Extensive experimental results on four real-world market datasets demonstrate the superiority of CMM over the current distillation method and RL-based market-making strategies. Tianhao Fu, Xinxin Xu 0006, Weichen Xu 0001, Ruilong Ren, Jian Cao 0002, Xixin Cao |
AAAI | 1 |
| 2026 | Latency-SLO-Aware Memory Offloading for Large Language Model InferenceabstractOffloading large language models (LLMs) states to host memory during inference promises to reduce operational costs by supporting larger models, longer prompts, and larger batch sizes. However, the design of existing memory offloading mechanisms does not take latency service-level objectives (SLOs) into consideration. As a result, they either lead to frequent SLO violations or underutilize host memory, thereby incurring economic loss and thus defeating the purpose of memory offloading. Chenxiang Ma, Zhisheng Ye 0002, Zehua Yang, Tianhao Fu, Jiaxun Han, Jie Zhang 0048, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Yong Li 0045, Diyu Zhou |
ICS | 5 |
| 2025 | Human-Inspired Situated Question Answering with Large Language ModelsabstractSituated Question Answering (SQA) involves reasoning the question based on the scene context from the questioner's perspective, enabling potential applications in human-centered embodied intelligence. Previous works handle this task by data-driven methods, generating answers in a closed-set or open-ended manner without explicit reasoning. Therefore, these works are often trapped in dataset bias, overfitting, and interpretability challenges. Inspired by the reasoning process of humans, we propose a Human-Inspired SQA model (Hi-QA), the first embodied agent that leverages human-like skills to understand the scene and perform interpretable reasoning. Specifically, Hi-QA leverages a specialized perception tool to gather information from 3D scenes for the reasoning module to reason questions with Large Language Models. The reflection module amends invalid reasoning and builds methodology memory to learn problem-solving methods. The distillation module distills prototypical information about objects, supporting the agent with supplementary knowledge. Experiments on ScanQA and SQA3D datasets validate that Hi-QA achieves state-of-the-art performance and highlight that human-inspired design significantly boosts the model's interpretability and generalizability. Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Ruilong Ren, Xing Zhang 0002 |
ICME | 4 |
| 2025 | Mutual information-driven self-supervised point cloud pre-training
Weichen Xu 0001, Tianhao Fu, Jian Cao 0002, Xinxin Xu 0006, Xixin Cao, Xing Zhang 0002 |
Knowl. Based Syst. | 2 |
| 2024 | J-MAE: Jigsaw Meets Masked Autoencoders in X-Ray Security InspectionabstractThe X-ray security inspection aims to identify any restricted items to protect public safety. Due to the lack of focus on unsupervised learning in this field, using pre-trained models on natural images leads to suboptimal results in downstream tasks. Previous works would lose the relative positional relationships during the pre-training process, which is detrimental for X-ray images that lack texture and rely on shape. In this paper, we propose the jigsaw style MAE (J-MAE) to preserve the relative position information by shuffling the position encoding of visible patches. This forces the network to perform semantic reasoning to understand the shape and composition of X-ray objects. Meanwhile, we propose the Incremental Shuffling Module (ISM) and Permute Predicting Module (PPM) to make the training process more stable and accelerate convergence. Our proposed method has consistently outperformed other methods on three downstream X-ray security inspection datasets. Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Awen Bai, Ruilong Ren, Zicong Hu, Xixin Cao, Xing Zhang 0002 |
ICASSP | 3 |
| 2024 | SweepMM: A High-Quality Multimodal Dataset for Sweeping Robots in Home Scenarios for Vision-Language ModelabstractEmbodied intelligence based on vision-language models aims to learn from interactions and derive general intelligence. However, existing generalized vision-language models cannot understand domain knowledge in home scenarios due to the lack of sweeping robot multimodal datasets. In this paper, we propose the first multimodal dataset for sweeping robots, called SweepMM. We create textual data such as room type, scene descriptions, and moving recommendations using various approaches including rule-based, manual-based, and off-the-shelf model-based methods. Based on this dataset, we fine-tune the first generative pretrained model for sweeping robots, called SweepGPM. This model enables human-robot dialogue and surpasses previous state-of-the-art methods by 0.8% in room type recognition, 0.4% in obstacle detection, and 8.0% in lost item search, demonstrating the potential of embodied intelligence in sweeping robots. Weichen Xu 0001, Xinxin Xu 0006, Tianhao Fu, Jian Cao 0002, Yuetian Huang, Xixin Cao, Xing Zhang 0002 |
ICASSP | 3 |
| 2024 | Empirical Research On Quantization For 3D Multi-Modal Vit ModelsabstractModel quantization finds success in simplifying model inference in practical applications. However, it predominantly focuses on CNNs and 2D ViT models, with limited attention given to quantizing 3D models. We pensively explore 3D model quantization challenges and discover similar numerical distributions of Softmax and LayerNorm between 3D and 2D models. Consequently, we apply the quantization algorithms FQ-ViT and I-ViT designed for 2D ViT models to 3D model quantization to address performance issues caused by uneven numerical distributions in Softmax and LayerNorm. Our research includes extensive experiments using transformer architectures and establishes benchmarks, demonstrating successful quantization of 3D multimodal model UNITR. Notably, our approach experiences a slight decrease compared to FP32 while outperforming other state-of-the-art models. For example, in the 3D object detection task on the nuScenes dataset, the 8 -bit UNITR (FQ-ViT) achieves impressive NDS and mAP scores of $73.0 \%$ and $70.0 \%$, surpassing the full precision BEVFusion model. Zicong Hu, Jian Cao 0002, Weichen Xu 0001, Ruilong Ren, Tianhao Fu, Xinxin Xu 0006, Xing Zhang 0002 |
ICIP | 5 |
| 2024 | Boosting 3D Visual Grounding by Object-Centric Referring Networkabstract3D visual grounding is tasked with locating a specific object within a 3D scene, as described by a given textual reference. This task is challenging because it requires (1) the accurate recognition of various objects in a 3D scene and (2) the understanding of spatial relations in the description. However, current studies encounter difficulties in situations where multiple similar objects are present or when the descriptions involve intricate and abstract relations. In this paper, a novel, simple, and efficient Object-Centric Referring network, namely 3D-OCR, is presented to take high-quality semantic representation and deep relation modeling into account. Specifically, an offline Fine-grained Semantic Enhancement (FSE) module is designed to reinforce the object-centric semantic awareness with fine-grained high-quality object semantic representations. To achieve superior object-centric relation awareness, we propose a Deep Relation Modeling (DRM) module with the explicit and implicit relation self-attention module, enriching object features with relational context. Moreover, we utilize a vision-language contrastive loss to further improve the matching process between point cloud and language. Comprehensive experiments conducted on the challenging ScanRefer and Nr3D datasets corroborate the exceptional performance of our method, with an increase of +1.47% on ScanRefer and +1.2% on Nr3D. Ruilong Ren, Jian Cao 0002, Weichen Xu 0001, Tianhao Fu, Yilei Dong, Xinxin Xu 0006, Zicong Hu, Xing Zhang 0002 |
IROS | 4 |
| 2024 | Point Cloud Reconstruction Is Insufficient to Learn 3D RepresentationsabstractThis paper revisits the development of generative self-supervised learning in 2D images and 3D point clouds in autonomous driving. In 2D images, the pretext task has evolved from low-level to high-level features. Inspired by this, through explore model analysis, we find that the gap in weight distribution between self-supervised learning and supervised learning is substantial when employing only low-level features as the pretext task in 3D point clouds. Low-level features represented by PoInt Cloud reconsTruction are insUfficient to learn 3D REpresentations (dubbed PICTURE). To advance the development of pretext tasks, we propose a unified generative self-supervised framework. Firstly, high-level features are demonstrated to exhibit semantic consistency with downstream tasks. We utilize the high-level features as an additional pretext task to enhance the understanding of semantic information during the pre-training. Next, we propose inter-class and intra-class discrimination-guided masking (I2Mask) based on the attributes of the high-level features, adaptively setting the masking ratio for each superclass. On Waymo and nuScenes datasets, we achieve 75.13% mAP and 72.69% mAPH for 3D object detection, 79.4% mIoU for 3D semantic segmentation, and 18.4% mIoU for occupancy prediction. Extensive experiments have demonstrated the effectiveness and necessity of high-level features. Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Ruilong Ren, Zicong Hu, Xixin Cao, Xing Zhang 0002 |
ACM Multimedia | 3 |
| 2023 | Overcoming the Seesaw in Monocular 3D Object Detection Via Language Knowledge TransferringabstractMonocular 3D object detection is a challenging problem in self-driving and computer vision communities. Previous works suffered from a severe seesaw phenomenon: multi-category learning was worse than single-category, and feature learning between categories inhibited each other. We reveal that the real culprit is the significant difference in depth distribution between categories. Confusing feature representations exacerbate depth estimation. In this paper, we propose Language Knowledge Transferring to introduce language information in monocular 3D object detection, termed as MonoLT. Multimodal language-Image guides networks learn more class-specific features, which reduces the pressure of depth estimation. Meanwhile, we propose the Polar Depth Aggregator to make the depth estimation less disturbed by the environment and other instances (especially different classes). Comprehensive experiments performed on the KITTI dataset prove the superiority of our proposed method. The code will be released soon. Weichen Xu 0001, Tianhao Fu |
ICASSP | 2 |
| 2022 | Boosting Dense Long-Tailed Object Detection from Data-Centric View
Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Hongyi Yao, Yuan Wang 0001 |
ACCV (3) | 3 |
| 2022 | Tear Up the Bubble Boom: Lessons Learned From a Deep Learning Research and Development ClusterabstractWith the proliferation of deep learning, there exists a strong need to efficiently operate GPU clusters for deep learning production in giant AI companies, as well as for research and development (R&D) in small-sized research institutes and universities. Existing works have performed thorough trace analysis on large-scale production-level clusters in giant companies, which discloses the characteristics of deep learning production jobs and motivates the design of scheduling frameworks. However, R&D clusters significantly differ from production-level clusters in both job properties and user behaviors, calling for a different scheduling mechanism. In this paper, we present a detailed workload characterization of an R&D cluster, CloudBrain-I, in a research institute, Peng Cheng Laboratory. After analyzing the fine-grained resource utilization, we discover a severe problem for R&D clusters, resource underutilization, which is especially important in R&D clusters while not characterised by existing works. We further investigate two specific underutilization phenomena and conclude several implications and lessons on R&D cluster scheduling. The traces will be open-sourced to motivate further studies in the community. Zehua Yang, Zhisheng Ye 0002, Tianhao Fu, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Tianwei Zhang 0004 |
ICCD | 3 |
| 2021 | Lifting the Veil of Frequency in Joint Segmentation and Depth EstimationabstractJoint learning of scene parsing and depth estimation remains a challenging task due to the rivalry between the two tasks. In this paper, we revisit the mutual enhancement for joint semantic segmentation and depth estimation. Inspired by the observation that the competition and cooperation could be reflected in the feature frequency components of different tasks, we propose a Frequency Aware Feature Enhancement (FAFE) network that can effectively enhance the reciprocal relationship whereas avoiding the competition. In FAFE, a frequency disentanglement module is proposed to fetch the favorable frequency component sets for each task and resolve the discordance between the two tasks. For task cooperation, we introduce a re-calibration unit to aggregate features of the two tasks, so as to complement task information with each other. Accordingly, the learning of each task can be boosted by the complementary task appropriately. Besides, a novel local-aware consistency loss function is proposed to impose on the predicted segmentation and depth so as to strengthen the cooperation. With the FAFE network and new local-aware consistency loss encapsulated into the multi-task learning network, the proposed approach achieves superior performance over previous state-of-the-art methods. Extensive experiments and ablation studies on multi-task datasets demonstrate the effectiveness of our proposed approach. Tianhao Fu, Xiaoqing Ye, Xiao Tan 0001, Fumin Shen, Errui Ding |
ACM Multimedia | 1 |