EDBT 2026 Demo / reviewers in the wild / expert
Yusong Hu
dblp:134/8808
· DBLP profile ↗
14ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0002-1495-2128ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAPO: Dynamic Teacher and Perturbed Answer Injection for Policy OptimizationabstractReinforcement learning (RL) has emerged as a powerful framework to improve the reasoning performance of large language models (LLMs), with approaches such as Group Relative Policy Optimization (GRPO) showing promising results. However, GRPO and its variants struggle with collapsed groups (i.e., all-correct or all-incorrect completions), leading to zero-variance rewards and ineffective gradient signals. Moreover, focusing solely on final answer correctness while ignoring the reasoning process, along with rigid length penalties, can hinder training stability and output quality. To address these issues, we introduce TAPO, a reinforcement learning framework that enhances optimization signals by modifying sampled completions within training groups. TAPO incorporates three core techniques: (1) Dynamic Teacher Injection (DTI), which selectively injects high-quality or adversarial examples to restore effective gradient signals in collapsed groups; (2) Perturbed Answer Injection (PAI), which makes partially correct completions to provide contrastive supervision separating reasoning correctness but wrong answer from the trajectories; and (3) InfoLen-Aware Reward Shaping, a fine-grained reward strategy that penalizes outputs based on both length and semantic redundancy, encouraging concise yet informative responses. Extensive experimental results demonstrate that TAPO significantly improves the mathematical reasoning capabilities of LLMs across multiple challenging benchmarks, outperforming the GRPO baseline by a substantial margin. Component-wise ablations further validate the contribution of each proposed technique. Maowei Jiang, Peter Bús, Moquan Chen, Quangao Liu, Ruiqi Li 0004, Pengyu Zeng, Ruikai Liu, Alan Liang, Yusong Hu, Zhiyong Dong |
AAAI | 13 |
| 2026 | FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge FlowabstractYusong Hu, Runmin Ma, Yue Fan, Jinxin Shi, Zongsheng Cao, Yuhao Zhou, Jiakang Yuan, Shuaiyu Zhang, Shiyang Feng, Xiangchao Yan, Shufei Zhang, Wenlong Zhang, Lei Bai, Bo Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yusong Hu, Runmin Ma, Jinxin Shi, Zongsheng Cao, Yuhao Zhou 0005, Jiakang Yuan, Shuaiyu Zhang, Shiyang Feng, Xiangchao Yan, Shufei Zhang, Lei Bai 0001, Bo Zhang 0069 |
ACL (1) | 1 |
| 2026 | Mixed reality and machine learning-guided cable robot framework (MMCR) for real-time prefabricated construction automationabstractDespite recent advances in automation, the AECO industry still faces inefficiencies, high costs, and safety risks. To address these, we present MMCR, a mixed-reality (MR) and machine-learning (ML) guided cable-robotic system for prefabricated construction. A HoloLens 2 MR interface delivers spatially anchored guidance and risk alerts, while a Unity ML agent enables autonomous navigation, obstacle avoidance, grasping, and placement. In real-world tests, the ML-enabled cable robot manipulated precast components in real time, reducing manual intervention and cognitive load. Compared with manual and MR-only baselines, MMCR achieved a 4 × reduction in positioning error, 2.8 × faster task completion, and 88.9% fewer operator interventions. These results indicate that MR-assisted, learning-based cable robotics can streamline onsite assembly and enhance safety, advancing practical pathways toward smart, efficient construction. Maowei Jiang, Tingtao Yu, Yusong Hu, Zhiyong Dong, Hongfei Ai, Peter Bús |
Adv. Eng. Informatics | 4 |
| 2025 | Docopilot: Improving Multimodal Models for Document-Level UnderstandingabstractDespite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lack of high-quality, document-level datasets. While current retrieval-augmented generation (RAG) methods offer partial solutions, they suffer from issues, such as fragmented retrieval contexts, multi-stage error accumulation, and extra time costs of retrieval. In this work, we present a high-quality document-level dataset, Doc-750K, designed to support in-depth understanding of multimodal documents. This dataset includes diverse document structures, extensive cross-page dependencies, and real question-answer pairs derived from the original documents. Building on the dataset, we develop a native multimodal model—Docopilot, which can accurately handle document-level dependencies without relying on RAG. Experiments demonstrate that Docopilot achieves superior coherence, accuracy, and efficiency in document understanding tasks and multi-turn interactions, setting a new baseline for document-level multimodal understanding. Data, code, and models are released at https://github.com/OpenGVLab/Docopilot. Yuchen Duan, Zhe Chen 0017, Yusong Hu, Weiyun Wang, Shenglong Ye, Botian Shi, Lewei Lu, Qibin Hou, Tong Lu 0002, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang |
CVPR | 3 |
| 2025 | KAC: Kolmogorov-Arnold Classifier for Continual LearningabstractContinual learning requires models to train continuously across consecutive tasks without forgetting. Most existing methods utilize linear classifiers, which struggle to maintain a stable classification space while learning new tasks. Inspired by the success of Kolmogorov-Arnold Networks (KAN) in preserving learning stability during simple continual regression tasks, we set out to explore their potential in more complex continual learning scenarios. In this paper, we introduce the Kolmogorov-Arnold Classifier (KAC), a novel classifier developed for continual learning based on the KAN structure. We delve into the impact of KAN’s spline functions and introduce Radial Basis Functions (RBF) for improved compatibility with continual learning. We replace linear classifiers with KAC in several recent approaches and conduct experiments across various continual learning benchmarks, all of which demonstrate performance improvements, highlighting the effectiveness and robustness of KAC in continual learning. The code is available at https://github.com/Ethanhuhuhu/KAC. Yusong Hu, Zichen Liang, Fei Yang 0004, Qibin Hou, Xialei Liu, Ming-Ming Cheng |
CVPR | 1 |
| 2025 | Achieving Plasticity-Stability Trade-Off in Continual Learning Through Adaptive Orthogonal ProjectionabstractCatastrophic forgetting is the crucial challenge for continual learning. One of the state-of-the-art approaches is the orthogonal projection, which aims to learn each task by updating model parameters in the direction orthogonal to the subspace spanned by the previous task input. Although such strict orthogonal weight constraints ensure no interference with tasks that have been learned to achieve model stability, they greatly sacrifice model plasticity. In this paper, we propose an adaptive balanced orthogonal projection (AdaBOP) method, to search for the optimal network parameter updating direction to address the plasticity-stability dilemma in continual learning. The proposed AdaBOP method can adaptively adjust its tendency towards plasticity-stability trade-off based on the layer-wise feature space correlations of the model between old and new tasks. To further improve the training efficiency, we also implement the AdaBOP method in the uncentered covariance matrix space of the previous tasks, and finally achieve a better stability-plasticity trade-off in continual learning efficiently. Experimental results greatly demonstrate the effectiveness of the proposed method, which achieves superior performances to state-of-the-art continual learning approaches. The code is available athttps://github.com/hyscn/AdaBOP. De Cheng, Yusong Hu, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Reformulating Classification as Image-Class Matching for Class Incremental LearningabstractClass incremental learning (CIL) sequentially increases the number of classes, which often leads to catastrophic forgetting when fine-tuning on new classes. Existing approaches typically employ linear classifiers and expand them to accommodate new classes. However, conducting conventional classification inherently introduces feature drift in the image space upon the introduction of new classifiers, potentially disrupting the established distributions, and resulting in forgetting. In this paper, we propose a novel insight to reformulate the conventional classification as image-class matching (ICM) to mitigate the disruption. ICM independently encodes the image and the category and allows for the sharing of a matching classifier across all tasks, effectively stabilizing the feature space during the CIL process. To apply ICM to CIL, we introduce the Binary Matching Classification (BMC) framework, which employs cross attention to encode the matching relationship between images and each category to predict matching scores. When learning new tasks, BMC only requires the addition of category inputs without any structural changes. Moreover, we present a series of strategies to enhance the adaptation of BMC to CIL. Through simple regularization, our BMC framework achieves outstanding performance on various benchmarks including CIFAR-100, ImageNet-100, and ImageNet-1000 datasets. Our code is available athttps://github.com/Ethanhuhuhu/BMC. Yusong Hu, Zichen Liang, Xialei Liu, Qibin Hou, Ming-Ming Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Enhancing Continual Semantic Segmentation via Uncertainty and Class Balance Re-WeightingabstractContinual Semantic Segmentation (CSS) primarily aims to continually learn new semantic segmentation categories while avoiding catastrophic forgetting. In semantic segmentation tasks, images can comprise both familiar old categories and novel unseen categories and they are treated as background in the incremental stage. Therefore, it is necessary to utilize the old model to generate pseudo-labels. However, the quality of these pseudo-labels significantly influences the model's forgetting of the old categories. Erroneous pseudo-labels can introduce harmful gradients, thus exacerbating model forgetting. In addition, the issue of class imbalance poses a significant challenge within the realm of CSS. Although traditional methods frequently diminish the emphasis placed on new classes to address this imbalance, we discover that the imbalance extends beyond the distinction between old and new classes. In this paper, we specifically address two previously overlooked problems in CSS: the impact of erroneous pseudo-labels on model forgetting and the confusion induced by class imbalance. We propose an Uncertainty and Class Balance Re-weighting approach (UCB) that assigns higher weights to pixels with pseudo-labels exhibiting lower uncertainty and to categories with smaller proportions during the training process. Our proposed approach enhances the impact of essential pixels during the continual learning process, thereby reducing model forgetting and dynamically balancing category weights based on the dataset. Our method is simple yet effective and can be applied to any method that uses pseudo-labels. Extensive experiments on the Pascal-VOC and ADE20K datasets demonstrate the efficacy of our approach in improving model performance across three state-of-the-art methods. The code will be available at https://github.com/JACK-Chen-2019/UCB. Zichen Liang, Yusong Hu, Fei Yang 0004, Xialei Liu |
IEEE Trans. Image Process. | 2 |
| 2024 | Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual LearningabstractContinual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis—one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task—the recent progress has achieved forget-free CL. Although addressing the CF issue well, such methods would encounter serious under-fitting in long-term CL, in which the learning process will continue for a long time and the number of new tasks involved will be much higher. To solve this problem, this paper partitions the network into three parts—with a new part for exploring the knowledge sharing between the old and new tasks. With the shared knowledge, this part of network can be learnt to simultaneously consolidate the old tasks and fit to the new task. To achieve this goal, we propose a task-aware Orthogonal Sparse Network (OSN), which contains shared knowledge induced network partition and sharpness-aware orthogonal sparse network learning. The former partitions the network to select shared parameters, while the latter guides the exploration of shared knowledge through shared parameters. Qualitative and quantitative analyses, show that the proposed OSN induces minimum to no interference with past tasks, i.e., approximately no forgetting, while greatly improves the model plasticity and capacity, and finally achieves the state-of-the-art performances. Yusong Hu, De Cheng, Dingwen Zhang, Nannan Wang 0001, Tongliang Liu, Xinbo Gao 0001 |
ICML | 1 |
| 2024 | Asymmetric Learned Image Compression Using Fast Residual Channel Attention
Yusong Hu, Cheolkon Jung, Yang Liu 0351 |
ICPR (26) | 1 |
| 2024 | A3R: Vision Language Pre-training by Attentive Alignment and Attentive Reconstruction
Yusong Hu, Ke Li 0015, Xialei Liu |
PRCV (5) | 1 |
| 2022 | Long-Tailed Class Incremental Learning
Xialei Liu, Yusong Hu, Andrew D. Bagdanov, Ke Li 0015, Ming-Ming Cheng |
ECCV (33) | 2 |
| 2022 | Analyzing Online Transaction Networks with Network MotifsabstractNetwork motif is a kind of frequently occurring subgraph that reflects local topology in graphs. Although network motif has been studied in graph analytics, e.g., social network and biological network, it is yet unclear whether network motif is useful for analyzing online transaction network that is generated in applications such as instant messaging and e-commerce. In this work, we analyze online transaction networks from the perspective of network motif. We define vertex features based on size-2 and size-3 motifs, and introduce motif-based centrality measurements. We further design motif-based vertex embedding that integrates weighted motif counts and centrality measurements. Afterward, we implement a distributed framework for motif detection in large-scale online transaction networks. To understand the effectiveness of motif for analyzing online transaction network, we study the statistical distribution of motifs in various kinds of graphs in Tencent and assess the benefit of motif-based embedding in a range of downstream graph analytical tasks. Empirical results show that our proposed method can efficiently find motifs in large-scale graphs, help interpretability, and benefit downstream tasks. Jiawei Jiang 0001, Yusong Hu, Xiaosen Li, Wen Ouyang, Zhitao Wang, Fangcheng Fu, Bin Cui 0001 |
KDD | 2 |
| 2015 | Memory-efficient discrete wavelet transform architecture based on wordlength optimizationabstractUnlike the existing designs that improve the memory efficiency by reducing the on-chip memory words, we propose a memory-efficient 2-D DWT architecture which improves the memory efficiency by reducing the wordlength of the on-chip memory. Based on our recently proposed memory-efficient 2-D DWT architecture, which achieves the highest memory efficiency among the existing designs, we analyze the dynamic range and optimize the required integer bit (IB) width of every internal signal for memory reduction. We construct an architecture-specific accuracy model for both 9/7 and 5/3 2-D DWT and optimize the factional bit (FB) width of every internal signal without sacrificing the precision of the output. Theoretical estimation shows a 17.5% reduction of the temporal memory regardless of input image size and throughput. In addition, the arithmetic resource is reduced. The synthesis results in 90-nm CMOS show a better area-delay product (ADP) of 25.1% over the best existing design. Yusong Hu, Ching-Chuen Jong |
ISCAS | 1 |