VLDB 2026 Research / reviewers in the wild / expert
Junyuan Shang
dblp:225/9295
· DBLP profile ↗
24ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-4301-750XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty-Aware Routing for Principled Alignment with MoE DynamicsabstractYilong Chen, Junyuan Shang, Yuchen Feng, Zhenyu Zhang, Naibin Gu, Ziqi Wang, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junyuan Shang, Zhenyu Zhang 0006, Naibin Gu, Tingwen Liu, Shuohuan Wang, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 2 |
| 2026 | Hierarchical Cross-Modality Interaction for Unified Video-Text Retrieval Modeling
Tianshi Xu, Zhengzheng Sun, Yizheng Hu, Junyuan Shang, Si Wu 0002 |
MMM (2) | 4 |
| 2025 | Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal ThinkingabstractYilong Chen, Junyuan Shang, Zhenyu Zhang, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junyuan Shang, Zhenyu Zhang 0006, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Haifeng Wang 0001 |
ACL (1) | 2 |
| 2025 | Mixture of Hidden-Dimensions: Not All Hidden-States' Dimensions are Needed in TransformerabstractTransformer models encounter inefficiency when scaling hidden dimensions due to the uniform expansion of parameters. When delving into the sparsity of hidden dimensions, we observe that only a small subset of dimensions are highly activated, where some dimensions are commonly activated across tokens, and some others uniquely activated for individual tokens. To leverage this, we propose MoHD (Mixture of Hidden Dimensions), a sparse architecture that combines shared sub-dimensions for common features and dynamically routes specialized sub-dimensions per token. To address the potential information loss from sparsity, we introduce activation scaling and group fusion mechanisms. MoHD efficiently expands hidden dimensions with minimal computational increases, outperforming vanilla Transformers in both parameter efficiency and task performance across 10 NLP tasks. MoHD achieves 1.7% higher performance with 50% fewer activatied parameters and 3.7% higher performance with 3$\times$ total parameters expansion at constant activated parameters cost. MoHD offers a new perspective for scaling the model, showcasing the potential of hidden dimension sparsity. Junyuan Shang, Zhenyu Zhang 0006, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003, Haifeng Wang 0001 |
ICML | 2 |
| 2025 | BiPC: Bidirectional Probability Calibration for Unsupervised Domain Adaption
Wenlve Zhou, Zhiheng Zhou 0001, Junyuan Shang, Chang Niu, Xiyuan Tao, Tianlei Wang |
Expert Syst. Appl. | 3 |
| 2024 | LEMON: Reviving Stronger and Smaller LMs from Larger LMs with Linear Parameter FusionabstractYilong Chen, Junyuan Shang, Zhenyu Zhang, Shiyao Cui, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junyuan Shang, Zhenyu Zhang 0006, Shiyao Cui, Tingwen Liu, Shuohuan Wang, Yu Sun 0029, Hua Wu 0003 |
ACL (1) | 2 |
| 2024 | NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference TimeabstractYilong Chen, Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang, Tingwen Liu, Shuohuan Wang, Yu Sun, Dianhai Yu, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang, Dianhai Yu, Hua Wu 0003 |
ACL (1) | 3 |
| 2024 | DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionabstractLarge language models (LLMs) with billions of parameters demonstrate impressive performance. However, the widely used Multi-Head Attention (MHA) in LLMs incurs substantial computational and memory costs during inference. While some efforts have optimized attention mechanisms by pruning heads or sharing parameters among heads, these methods often lead to performance degradation or necessitate substantial continued pre-training costs to restore performance. Based on the analysis of attention redundancy, we design a Decoupled-Head Attention (DHA) mechanism. DHA adaptively configures group sharing for key heads and value heads across various layers, achieving a better balance between performance and efficiency. Inspired by the observation of clustering similar heads, we propose to progressively transform the MHA checkpoint into the DHA model through linear fusion of similar head parameters step by step, retaining the parametric knowledge of the MHA checkpoint. We construct DHA models by transforming various scales of MHA checkpoints given target head budgets. Our experiments show that DHA remarkably requires a mere 0.25\% of the original model's pre-training budgets to achieve 96.1\% of performance while saving 75\% of KV cache. Compared to Group-Query Attention (GQA), DHA achieves a 5$\times$ training acceleration, a maximum of 13.93\% performance improvement under 0.01\% pre-training budget, and 5\% relative improvement under 0.05\% pre-training budget. Linhao Zhang, Junyuan Shang, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang |
NeurIPS | 3 |
| 2024 | Superclass-aware visual feature disentangling for generalized zero-shot learning
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junmei Yang |
Expert Syst. Appl. | 2 |
| 2024 | Consistent representation joint adaptive adjustment for incremental zero-shot learning
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junmei Yang |
Neurocomputing | 2 |
| 2024 | Generalized zero-shot action recognition through reservation-based gate and semantic-enhanced contrastive learning
Junyuan Shang, Chang Niu, Xiyuan Tao, Zhiheng Zhou 0001, Junmei Yang |
Knowl. Based Syst. | 1 |
| 2022 | Asymmetric Adversarial-based Feature Disentanglement Learning for Cross-Database Micro-Expression RecognitionabstractRecently, micro-expression recognition (MER) has gained tremendous progress. However, most methods are based on individual-database micro-expression recognition and are difficult to generalize into complicated scenarios. Therefore, cross-database micro-expression recognition (CDMER) has drawn growing attention due to its robustness and generalizability. In this paper, we propose a novel CDMER algorithm with asymmetric adversarial-based feature disentanglement learning, which implements the disentanglement of domain features and emotion features aiming to learn domain-invariant and discriminative representation. Furthermore, to facilitate the feature disentanglement learning, a Domain Information Filtering (DF) module is designed to filter out the domain component of the micro-expression features (emotion feature). Extensive experiments on the SMIC and CASME II databases have shown that our proposed method outperforms the state-of-the-art method and has superior performance against excessive domain discrepancies. Zhiheng Zhou 0001, Junyuan Shang |
ACM Multimedia | 3 |
| 2022 | Few-shot domain adaptation through compensation-guided progressive alignment and bias reduction
Junyuan Shang, Chang Niu, Junchu Huang, Zhiheng Zhou 0001, Junmei Yang |
Appl. Intell. | 1 |
| 2022 | Unbiased feature generating for generalized zero-shot learning
Chang Niu, Junyuan Shang, Junchu Huang, Junmei Yang, Yuting Song, Zhiheng Zhou 0001, Guoxu Zhou |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Label-guided heterogeneous domain adaptation
Zhiheng Zhou 0001, Chang Niu, Junyuan Shang |
Multim. Tools Appl. | 4 |
| 2021 | ERNIE-Doc: A Retrospective Long-Document Modeling TransformerabstractSiYu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun 0004, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Asymmetric alignment joint consistent regularization for multi-source domain adaptation
Junyuan Shang, Chang Niu, Zhiheng Zhou 0001, Junchu Huang, Zhiwei Yang 0010, Xiangwei Li |
Multim. Tools Appl. | 1 |
| 2020 | Common-specific feature learning for multi-source domain adaptationabstractMulti‐source domain adaptation (MDA) aims to leverage knowledge from multiple source domains to improve the classification performance on target domains. Different degrees of distribution discrepancies between every two domains pose a huge challenge to MDA tasks. Most works focus on extracting features shared by all domains, which is critical but not enough to reduce distribution discrepancies. In this paper, we propose a method named as common‐specific feature learning (CSFL). Constituting a framework of feature learning, CSFL explores a subspace where the combination of common and specific features makes learned representations comprehensive. Based on this framework, we conduct a metric learning method for learning a discriminative feature representation. Considering redundant information caused by source domains is likely to hurt the performance, we impose an effective low‐rank constraint to remove the redundant information. Further, we adopt structure consistent constraint to preserve the local structure in each domain. CSFL has obtained about 1–5% improvement of mean accuracy, compared to the state‐of‐the‐art shallow methods. Further, compared with 90.2% and 89.4% of the best baseline deep method, CSFL achieves mean accuracy of 90.8% and 89.7% on the Office‐31 and ImageCLEF‐DA datasets respectively. The encouraging results validate the effectiveness of our method. Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junchu Huang, Tianlei Wang, Xiangwei Li |
IET Image Process. | 2 |
| 2020 | Knowledge-shot learning: An interpretable deep model for classifying imbalanced electrocardiography data
Yen-hsiu Chou, Shenda Hong, Junyuan Shang, Moxian Song, Hongyan Li 0002 |
Neurocomputing | 4 |
| 2020 | Heterogeneous domain adaptation with label and structural consistency
Junchu Huang, Zhiheng Zhou 0001, Junyuan Shang, Chang Niu |
Multim. Tools Appl. | 3 |
| 2019 | GAMENet: Graph Augmented MEmory Networks for Recommending Medication CombinationabstractRecent progress in deep learning is revolutionizing the healthcare domain including providing solutions to medication recommendations, especially recommending medication combination for patients with complex health conditions. Existing approaches either do not customize based on patient health history, or ignore existing knowledge on drug-drug interactions (DDI) that might lead to adverse outcomes. To fill this gap, we propose the Graph Augmented Memory Networks (GAMENet), which integrates the drug-drug interactions knowledge graph by a memory module implemented as a graph convolutional networks, and models longitudinal patient records as the query. It is trained end-to-end to provide safe and personalized recommendation of medication combination. We demonstrate the effectiveness and safety of GAMENet by comparing with several state-of-the-art methods on real EHR data. GAMENet outperformed all baselines in all effectiveness measures, and also achieved 3.60% DDI rate reduction from existing EHR data. Junyuan Shang, Cao Xiao, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001 |
AAAI | 1 |
| 2019 | Pre-training of Graph Augmented Transformers for Medication RecommendationabstractMedication recommendation is an important healthcare application. It is commonly formulated as a temporal prediction task. Hence, most existing works only utilize longitudinal electronic health records (EHRs) from a small number of patients with multiple visits ignoring a large number of patients with a single visit (selection bias). Moreover, important hierarchical knowledge such as diagnosis hierarchy is not leveraged in the representation learning process. Despite the success of deep learning techniques in computational phenotyping, most previous approaches have two limitations: task-oriented representation and ignoring hierarchies of medical codes. To address these challenges, we propose G-BERT, a new model to combine the power of Graph Neural Networks (GNNs) and BERT (Bidirectional Encoder Representations from Transformers) for medical code representation and medication recommendation. We use GNNs to represent the internal hierarchical structures of medical codes. Then we integrate the GNN representation into a transformer-based visit encoder and pre-train it on EHR data from patients only with a single visit. The pre-trained visit encoder and representation are then fine-tuned for downstream predictive tasks on longitudinal EHRs from patients with multiple visits. G-BERT is the first to bring the language model pre-training schema into the healthcare domain and it achieved state-of-the-art performance on the medication recommendation task. Junyuan Shang, Tengfei Ma 0001, Cao Xiao, Jimeng Sun 0001 |
IJCAI | 1 |
| 2019 | K-margin-based Residual-Convolution-Recurrent Neural Network for Atrial Fibrillation DetectionabstractAtrial Fibrillation (AF) is an abnormal heart rhythm which can trigger cardiac arrest and sudden death. Nevertheless, its interpretation is mostly done by medical experts due to high error rates of computerized interpretation. One study found that only about 66% of AF were correctly recognized from noisy ECGs. This is in part due to insufficient training data, class skewness, as well as semantical ambiguities caused by noisy segments in an ECG record. In this paper, we propose a K-margin-based Residual-Convolution-Recurrent neural network (K-margin-based RCR-net) for AF detection from noisy ECGs. In detail, a skewness-driven dynamic augmentation method is employed to handle the problems of data inadequacy and class imbalance. A novel RCR-net is proposed to automatically extract both long-term rhythm-level and local heartbeat-level characters. Finally, we present a K-margin-based diagnosis model to automatically focus on the most important parts of an ECG record and handle noise by naturally exploiting expected consistency among the segments associated for each record. The experimental results demonstrate that the proposed method with 0.8125 F1NAOP score outperforms all state-of-the-art deep learning methods for AF detection task by 6.8%. Shenda Hong, Junyuan Shang, Hongyan Li 0002, Junqing Xie |
IJCAI | 3 |
| 2018 | Knowledge Guided Multi-instance Multi-label Learning via Neural Networks in Medicines PredictionabstractPredicting medicines for patients with co-morbidity has long been recognized as a hard task due to complex dependencies between diseases and medicines. Efforts have been made recently to build high-order dependency between diseases and medicines by extracting knowledge from electronic health records (EHR). But current works failed to utilize additional knowledge and ignored the data skewness problem which lead to sub-optimal combination of medicines. In this paper, we formulate the medicines prediction task in multi-instance multi-label learning framework considering the multi-diagnoses as input instances and multi-medicines as output labels. We propose a knowledge-guided multi-instance multi-label networks called \mname where two types of additional knowledge are incorporated into a RNN encoder-decoder model. The utilization of structural knowledge like clinical ontology provides a way to learn better representation called tree embedding by utilizing the ancestors’ information. Contextual knowledge is a global summarization of input instances which is informative for personal prediction. Experiments are conducted on a real world clinical dataset which showed the necessity to combine both contextual and structural knowledge and the \mname performs better than baselines up to 4+% in terms of Jaccard similarity score. Junyuan Shang, Shenda Hong, Hongyan Li 0002 |
ACML | 1 |