Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sihao Lin

dblp:254/5807 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0004-3235-373XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Efficient and distributed learning · 46% Image recognition and object detection · 21% Deep learning architectures and training · 11%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
3.642026
Distilling Object Detectors via Monte Carlo Dropout · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Entropy-Guided Condensing for Vision Transformer · Int. J. Comput. Vis. 2026
Let LLM Tell What to Prune and How Much to Prune · ICML 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
2.832026
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Entropy-Guided Condensing for Vision Transformer · Int. J. Comput. Vis. 2026
MLP Can Be a Good Transformer Learner · CVPR 2024
Computer vision › Image recognition and object detection
object detection
2.342026
Distilling Object Detectors via Monte Carlo Dropout · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Context Matters: Distilling Knowledge Graph for Enhanced Object Detection · IEEE Trans. Multim. 2024
Semi-Supervised Human Detection via Region Proposal Networks Aided by Verification · IEEE Trans. Image Process. 2020
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
2.132026
Distilling Object Detectors via Monte Carlo Dropout · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Knowledge Distillation via the Target-aware Transformer · CVPR 2022
Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation · ICCV 2021
Machine learning › Efficient and distributed learning
efficient training
1.012026
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Image recognition and object detection › object detection
knowledge distillation for detection
1.012026
Distilling Object Detectors via Monte Carlo Dropout · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Efficient and distributed learning › efficient training
progressive learning
1.012026
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Efficient and distributed learning › token reduction
token condensation
1.012026
Entropy-Guided Condensing for Vision Transformer · Int. J. Comput. Vis. 2026
Machine learning › Efficient and distributed learning
token reduction
1.012026
Entropy-Guided Condensing for Vision Transformer · Int. J. Comput. Vis. 2026
Computer vision › Image recognition and object detection
pedestrian detection
1.022022
Unreliable-to-Reliable Instance Translation for Semi-Supervised Pedestrian Detection · IEEE Trans. Multim. 2022
Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual Reinforcement · ICCV 2019
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble bootstrapping
0.912025
BossNAS Family: Block-Wisely Self-Supervised Neural Architecture Search · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning › model compression
large language model compression
0.912025
Let LLM Tell What to Prune and How Much to Prune · ICML 2025
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.912025
BossNAS Family: Block-Wisely Self-Supervised Neural Architecture Search · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.912025
Let LLM Tell What to Prune and How Much to Prune · ICML 2025
Computer vision › Image recognition and object detection › object detection
contextual reasoning
0.812024
Context Matters: Distilling Knowledge Graph for Enhanced Object Detection · IEEE Trans. Multim. 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.812024
Context Matters: Distilling Knowledge Graph for Enhanced Object Detection · IEEE Trans. Multim. 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning
0.812024
Making Large Language Models Better Planners with Reasoning-Decision Alignment · ECCV (36) 2024
Computer vision › 3D vision
3d object detection
0.712023
FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration · ICCV 2023
Robotics › Autonomous driving › perception
3d perception
0.712023
FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration · ICCV 2023
Computer vision › 3D vision › 3d object detection
multimodal 3d object detection
0.712023
FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration · ICCV 2023
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.612022
Unreliable-to-Reliable Instance Translation for Semi-Supervised Pedestrian Detection · IEEE Trans. Multim. 2022
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
feature distillation
0.512021
Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation · ICCV 2021
Computer vision › Image recognition and object detection › object detection
object proposal generation
0.412020
Semi-Supervised Human Detection via Region Proposal Networks Aided by Verification · IEEE Trans. Image Process. 2020
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection
0.412020
Semi-Supervised Human Detection via Region Proposal Networks Aided by Verification · IEEE Trans. Image Process. 2020
Machine learning › Generative modeling › generative adversarial network › conditional GAN
class-conditional GAN
0.412019
Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual Reinforcement · ICCV 2019
Machine learning › Generative modeling
diffusion model
0.312026
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Trustworthy machine learning
uncertainty estimation
0.312026
Distilling Object Detectors via Monte Carlo Dropout · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Efficient and distributed learning
parameter sharing
0.312025
BossNAS Family: Block-Wisely Self-Supervised Neural Architecture Search · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning
throughput optimization
0.212024
MLP Can Be a Good Transformer Learner · CVPR 2024
Computer vision › Segmentation and scene understanding › image segmentation
map segmentation
0.212023
FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration · ICCV 2023

Methods — techniques the papers use, named apart from their topics

transformer · 1.3vision transformer · 1.0one-shot growth schedule search · 1.0monte carlo dropout · 1.0momentum growth · 1.0information theory · 1.0entropy-guided condensing · 1.0automated progressive learning · 1.0importance quantification · 0.9dynamic pruning ratio · 0.9
YearPublicationVenuePosition
2026 Entropy-Guided Condensing for Vision Transformer
Sihao Lin, Pumeng Lyu, Dongrui Liu, Zhihui Li 0001, Wenguan Wang, Xiaojun Chang, Yuhui Zheng
Int. J. Comput. Vis.1
2026 Efficient Training of Large Vision Models via Advanced Automated Progressive Learning
abstract
The rapid advancements in Large Vision Models (LVMs), such as Vision Transformers (ViTs), diffusion models, and visual autoregressive models, have led to an increasing demand for computational resources, resulting in substantial financial and environmental costs. This growing challenge highlights the necessity of developing efficient training methods for LVMs. Progressive learning, a training strategy in which model capacity gradually increases during training, has shown promise in addressing these challenges. In this paper, we take a practical step toward the efficient training of LVMs by automating progressive learning. We focus first on the pre-training of LVMs, using ViTs as a case study. We propose AutoProg-One, an automated progressive learning scheme featuring momentum growth (MoGrow) and the one-shot growth schedule search. Additionally, we extend our approach beyond pre-training to address the transfer learning and fine-tuning of LVMs. We also expand the scope of AutoProg to encompass a wider range of LVMs, including diffusion models and visual autoregressive model. First, we introduce AutoProg-Zero, by enhancing the AutoProg framework with a novel zero-shot automated progressive learning method, eliminating the need for one-shot supernet training. Second, we introduce a novel Unique Stage Identifier (SID) scheme to bridge the gap during network growth. These innovations, integrated with the core principles of AutoProg, offer a comprehensive solution for efficient training across various LVM scenarios. Extensive experiments show that AutoProg accelerates ViT pre-training by up to 1.85 ×on ImageNet and accelerates the fine-tuning of diffusion models, and visual autoregressive model by up to 2.86 × and 1.89 ×, with comparable or even better performance. This work provides a robust and scalable approach to efficient training of LVMs, with potential applications in a wide range of vision tasks.
Sihao Lin, Zongxin Yang, Junwei Liang 0001, Xiaodan Liang, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Distilling Object Detectors via Monte Carlo Dropout
abstract
Knowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%.
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Let LLM Tell What to Prune and How Much to Prune
abstract
Large language models (LLMs) have revolutionized various AI applications. However, their billions of parameters pose significant challenges for practical deployment. Structured pruning is a hardware-friendly compression technique and receives widespread attention. Nonetheless, existing literature typically targets a single structure of LLMs. We observe that the structure units of LLMs differ in terms of inference cost and functionality. Therefore, pruning a single structure unit in isolation often results in an imbalance between performance and efficiency. In addition, previous works mainly employ a prescribed pruning ratio. Since the significance of LLM modules may vary, it is ideal to distribute the pruning load to a specific structure unit according to its role within LLMs. To address the two issues, we propose a pruning method that targets multiple LLM modules with dynamic pruning ratios. Specifically, we find the intrinsic properties of LLMs can guide us to determine the importance of each module and thus distribute the pruning load on demand, i.e., what to prune and how much to prune. This is achieved by quantifying the complex interactions within LLMs. Extensive experiments on multiple benchmarks and LLM variants demonstrate that our method effectively balances the trade-off between efficiency and performance.
Mingzhe Yang, Sihao Lin, Xiaojun Chang
ICML2
2025 BossNAS Family: Block-Wisely Self-Supervised Neural Architecture Search
abstract
Recent advances in hand-crafted neural architectures for visual recognition underscore the pressing need to explore architecture designs comprising diverse building blocks. Concurrently, neural architecture search (NAS) methods have gained traction as a means to alleviate human efforts. Nevertheless, the question of whether NAS methods can efficiently and effectively manage diversified search spaces featuring disparate candidates, such as Convolutional Neural Networks (CNNs) and transformers, remains an open question. In this work, we introduce a novel unsupervised NAS approach called BossNAS (Block-wisely Self-supervised Neural Architecture Search), which aims to address the problem of inaccurate predictive architecture ranking caused by a large weight-sharing space while mitigating potential ranking issue caused by biased supervision. To achieve this, we factorize the search space into blocks and introduce a novel self-supervised training scheme called Ensemble Bootstrapping, to train each block separately in an unsupervised manner. In the search phase, we propose an unsupervised Population-Centric Search, optimizing the candidate architecture towards the population center. Additionally, we enhance our NAS method by integrating masked image modeling and present BossNAS++ to overcome the lack of dense supervision in our block-wise self-supervised NAS. In BossNAS++, we introduce the training technique named Masked Ensemble Bootstrapping for block-wise supernet, accompanied by a Masked Population-Centric Search scheme to promote fairer architecture selection. Our family of models, discovered through BossNAS and BossNAS++, delivers impressive results across various search spaces and datasets. Our transformer model discovered by BossNAS++ attains a remarkable accuracy of 83.2% on ImageNet with only 10.5B MAdds, surpassing DeiT-B by 1.4% while maintaining a lower computation cost. Moreover, our approach excels in architecture rating accuracy, achieving Spearman correlations of 0.78 and 0.76 on the canonical MBConv search space with ImageNet and the NATS-Bench size search space with CIFAR-100, respectively, outperforming state-of-the-art NAS methods.
Sihao Lin, Tang Tao, Guangrun Wang, Mingjie Li 0006, Xiaodan Liang, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 MLP Can Be a Good Transformer Learner
abstract
Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and require same memory costs. This paper introduces a novel strategy that simplifies vision transformers and reduces computational load through the selective removal of non-essential attention layers, guided by entropy considerations. We identify that regarding the attention layer in bottom blocks, their subsequent MLP layers, i.e. two feed-forward layers, can elicit the same entropy quantity. Meanwhile, the accompanied MLPs are under-exploited since they exhibit smaller feature entropy compared to those MLPs in the top blocks. Therefore, we propose to integrate the uninformative attention layers into their subsequent counterparts by degenerating them into identical mapping, yielding only MLP in certain transformer blocks. Experimental results on ImageNet-1k show that the proposed method can remove 40% attention layer of DeiT-B, improving throughput and memory bound without performance compromise.
Sihao Lin, Pumeng Lyu, Dongrui Liu, Xiaodan Liang, Andy Song, Xiaojun Chang
CVPR1
2024 Making Large Language Models Better Planners with Reasoning-Decision Alignment
Shaoxiang Chen 0001, Sihao Lin, Zequn Jie, Lin Ma 0002, Guangrun Wang, Xiaodan Liang
ECCV (36)4
2024 Context Matters: Distilling Knowledge Graph for Enhanced Object Detection
abstract
The human visual system is capable of not only recognizing individual objects but also comprehending the contextual relationship between them in real-world scenarios, making it highly advantageous for object detection. However, in practical applications, such contextual information is often not available. Previous attempts to compensate for this by utilizing cross-modal data such as language and statistics to obtain contextual priors have been deemed sub-optimal due to a semantic gap. To overcome this challenge, we present a seamless integration of context into an object detector through Knowledge Distillation. Our approach intuitively represents context as a knowledge graph, describing the relative location and semantic relevance of different visual concepts. Leveraging recent advancements in graph representation learning with Transformer, we exploit the contextual information among objects using edge encoding and graph attention. Specifically, each image region propagates and aggregates the representation from its highly similar neighbors to form the knowledge graph in the Transformer encoder. Extensive experiments and a thorough ablation study conducted on challenging benchmarks MS-COCO, Pascal VOC and LVIS demonstrate the superiority of our method.
Aijia Yang, Sihao Lin, Chung-Hsing Yeh, Minglei Shu, Yi Yang 0001, Xiaojun Chang
IEEE Trans. Multim.2
2023 FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration
abstract
Multi-modality fusion and multi-task learning are becoming trendy in 3D autonomous driving scenario, considering robust prediction and computation budget. However, naively extending the existing framework to the domain of multi-modality multi-task learning remains ineffective and even poisonous due to the notorious modality bias and task conflict. Previous works manually coordinate the learning framework with empirical knowledge, which may lead to sub-optima. To mitigate the issue, we propose a novel yet simple multi-level gradient calibration learning framework across tasks and modalities during optimization. Specifically, the gradients, produced by the task heads and used to update the shared backbone, will be calibrated at the backbone’s last layer to alleviate the task conflict. Before the calibrated gradients are further propagated to the modality branches of the backbone, their magnitudes will be calibrated again to the same level, ensuring the downstream tasks pay balanced attention to different modalities. Experiments on large-scale benchmark nuScenes demonstrate the effectiveness of the proposed method, e.g., an absolute 14.4% mIoU improvement on map segmentation and 1.4% mAP improvement on 3D detection, advancing the application of 3D autonomous driving in the domain of multi-modality fusion and multi-task learning. We also discuss the links between modalities and tasks.
Sihao Lin, Guiyu Liu, Mukun Luo, Chaoqiang Ye, Hang Xu 0004, Xiaojun Chang, Xiaodan Liang
ICCV2
2022 Knowledge Distillation via the Target-aware Transformer
abstract
Knowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a one-to-one spatial matching fashion. However, people tend to overlook the fact that, due to the architecture differences, the semantic information on the same spatial location usually vary. This greatly undermines the underlying assumption of the one-to-one distillation approach. To this end, we propose a novel one-to-all spatial matching knowledge distillation approach. Specifically, we allow each pixel of the teacher feature to be distilled to all spatial locations of the student features given its similarity, which is generated from a target-aware transformer. Our approach surpasses the state-of-the-art methods by a significant margin on various computer vision benchmarks, such as ImageNet, Pascal VOC and COCOStuff10k. Code is available at https://github.com/sihaoevery/TaT.
Sihao Lin, Hongwei Xie, Bing Wang 0013, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang
CVPR1
2022 Unreliable-to-Reliable Instance Translation for Semi-Supervised Pedestrian Detection
abstract
Generating realistic pedestrian instances in a semi-supervised setting is promising but challenging due to the limited labeled data. We propose an unreliable-to-reliable instance translation model (Un2Reliab) conditioned on unreliable instances which poorly align with pedestrians. Un2Reliab mainly consists of an encoder-decoder-like generative network and a discriminative network, which are jointly trained in a minimax game. We adopt regularization to ensure that the synthesized instances are semantically similar to the corresponding ground truth. Furthermore, to preserve the identities of persons, we propose another regularization to ensure that the synthesized instances associated with the same person should be consistent in appearance. As a result, Un2Reliab learns to restore the missing parts of the original instances. As a side benefit, the synthesized instances are brought into better alignment. Inclusion of the synthesized data improves both the diversity and quality of training data, which eventually leads to better generalization performance. Extensive experiments indicate that Un2Reliab is able to synthesize high-fidelity pedestrian instances and improve the previous state-of-the-art results on multiple semi-supervised pedestrian detection benchmarks.
Sihao Lin, Si Wu 0002, Yong Xu 0007, Hau-San Wong
IEEE Trans. Multim.1
2021 Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation
abstract
Knowledge Distillation has shown very promising ability in transferring learned representation from the larger model (teacher) to the smaller one (student). Despite many efforts, prior methods ignore the important role of retaining inter-channel correlation of features, leading to the lack of capturing intrinsic distribution of the feature space and sufficient diversity properties of features in the teacher network. To solve the issue, we propose the novel Inter-Channel Correlation for Knowledge Distillation (ICKD), with which the diversity and homology of the feature space of the student network can align with that of the teacher network. The correlation between these two channels is interpreted as diversity if they are irrelevant to each other, otherwise homology. Then the student is required to mimic the correlation within its own embedding space. In addition, we introduce the grid-level inter-channel correlation, making it capable of dense prediction tasks. Extensive experiments on two vision tasks, including ImageNet classification and Pascal VOC segmentation, demonstrate the superiority of our ICKD, which consistently outperforms many existing methods, advancing the state-of-the-art in the fields of Knowledge Distillation. To our knowledge, we are the first method based on knowledge distillation boosts ResNet18 beyond 72% Top-1 accuracy on ImageNet classification. Code is available at: https://github.com/ADLab-AutoDrive/ICKD.
Li Liu 0069, Qingle Huang, Sihao Lin, Hongwei Xie, Bing Wang 0013, Xiaojun Chang, Xiaodan Liang
ICCV3
2020 Semi-Supervised Human Detection via Region Proposal Networks Aided by Verification
abstract
In this paper, we explore how to leverage readily available unlabeled data to improve semi-supervised human detection performance. For this purpose, we specifically modify the region proposal network (RPN) for learning on a partially labeled dataset. Based on commonly observed false positive types, a verification module is developed to assess foreground human objects in the candidate regions to provide an important cue for filtering the RPN's proposals. The remaining proposals with high confidence scores are then used as pseudo annotations for re-training our detection model. To reduce the risk of error propagation in the training process, we adopt a self-paced training strategy to progressively include more pseudo annotations generated by the previous model over multiple training rounds. The resulting detector re-trained on the augmented data can be expected to have better detection performance. The effectiveness of the main components of this framework is verified through extensive experiments, and the proposed approach achieves state-of-the-art detection results on multiple scene-specific human detection benchmarks in the semi-supervised setting.
Si Wu 0002, Shiyao Lei, Sihao Lin, Rui Li 0045, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Image Process.4
2019 Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual Reinforcement
abstract
We propose a GAN-based scene-specific instance synthesis and classification model for semi-supervised pedestrian detection. Instead of collecting unreliable detections from unlabeled data, we adopt a class-conditional GAN for synthesizing pedestrian instances to alleviate the problem of insufficient labeled data. With the help of a base detector, we integrate pedestrian instance synthesis and detection by including a post-refinement classifier (PRC) into a minimax game. A generator and the PRC can mutually reinforce each other by synthesizing high-fidelity pedestrian instances and providing more accurate categorical information. Both of them compete with a class-conditional discriminator and a class-specific discriminator, such that the four fundamental networks in our model can be jointly trained. In our experiments, we validate that the proposed model significantly improves the performance of the base detector and achieves state-of-the-art results on multiple benchmarks. As shown in Figure 1, the result indicates the possibility of using inexpensively synthesized instances for improving semi-supervised detection models.
Si Wu 0002, Sihao Lin, Mohamed Azzam, Hau-San Wong
ICCV2