VLDB 2026 Research / reviewers in the wild / expert
Weigao Sun
dblp:227/2932
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-2551-924XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Deep learning architectures and training · 48% Language models and text generation · 32% Efficient and distributed learning · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 78% Parallel and multicore computing · 12% High-performance computing · 10% |
Topics — the 26 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
1.8 | 2 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention · ICML 2024 |
Machine learning › Deep learning architectures and training
recurrent neural network |
1.7 | 2 | 2025 | Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025 Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025 |
Distributed systems › distributed machine learning
distributed training |
1.0 | 2 | 2025 | CO2: Efficient Distributed Training with Full Communication-Computation Overlap · ICLR 2024 Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters · AAAI 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
hybrid attention |
1.0 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Natural language and speech › Language models and text generation › language modeling
language model architecture |
1.0 | 1 | 2026 | Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation |
0.9 | 1 | 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.9 | 1 | 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025 |
Natural language and speech › Language models and text generation › text generation › surface realization
linearization |
0.9 | 1 | 2025 | Liger: Linearizing Large Language Models to Gated Recurrent Structures · ICML 2025 |
Machine learning › Efficient and distributed learning › efficient training
long-context training |
0.9 | 1 | 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters · AAAI 2025 |
Machine learning › Efficient and distributed learning › distributed training
data parallel training |
0.8 | 1 | 2024 | CO2: Efficient Distributed Training with Full Communication-Computation Overlap · ICLR 2024 |
Natural language and speech › Language models and text generation › efficient language model
efficient language model architectures |
0.8 | 1 | 2024 | Scaling Laws for Linear Complexity Language Models · EMNLP 2024 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.8 | 1 | 2024 | Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention · ICML 2024 |
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention |
0.8 | 1 | 2024 | Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention · ICML 2024 |
Machine learning › Deep learning architectures and training
scaling laws |
0.8 | 1 | 2024 | Scaling Laws for Linear Complexity Language Models · EMNLP 2024 |
Distributed systems › communication optimization
communication-computation overlap |
0.8 | 1 | 2024 | CO2: Efficient Distributed Training with Full Communication-Computation Overlap · ICLR 2024 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.4 | 1 | 2020 | pbSGD: Powered Stochastic Gradient Descent Methods for Accelerated Non-Convex Optimization · IJCAI 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
efficient sequence modeling |
0.3 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling |
0.3 | 1 | 2026 | Native Hybrid Attention for Efficient Sequence Modeling · ACL (1) 2026 |
Machine learning › Transfer learning and domain adaptation › model adaptation
model specialization |
0.3 | 1 | 2026 | Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism · ACL (1) 2026 |
Parallel and multicore computing
pipeline parallelism |
0.3 | 1 | 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters · AAAI 2025 |
Mathematical optimization
control theory |
0.3 | 1 | 2025 | Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2024 | Scaling Laws for Linear Complexity Language Models · EMNLP 2024 |
High-performance computing
large-scale training |
0.2 | 1 | 2024 | CO2: Efficient Distributed Training with Full Communication-Computation Overlap · ICLR 2024 |
Machine learning › Efficient and distributed learning › efficient training
training acceleration |
0.1 | 1 | 2020 | pbSGD: Powered Stochastic Gradient Descent Methods for Accelerated Non-Convex Optimization · IJCAI 2020 |
Methods — techniques the papers use, named apart from their topics
linear attention · 1.8sequence accumulation · 1.7pipeline parallelism · 1.7task-aware memory · 1.0softmax attention · 1.0sliding window attention · 1.0state feedback · 0.9output feedback · 0.9low-rank adaptation · 0.9hybrid attention · 0.9hallucination detection · 0.9delta learning rule · 0.9chunk-wise parallel kernel · 0.9automatic annotation pipeline · 0.9staleness gap penalty · 0.8outer momentum clipping · 0.8local updating · 0.8asynchronous communication · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Native Hybrid Attention for Efficient Sequence ModelingabstractTransformers excel at sequence modeling but face quadratic complexity, while linear attention offers improved efficiency but often compromises recall accuracy over long contexts.In this work, we introduce Native Hybrid Attention (NHA), a novel hybrid architecture of linear and full attention that integrates both intra & inter-layer hybridization into a unified layer design.NHA maintains longterm context in key-value slots updated by a linear RNN, and augments them with shortterm tokens from a sliding window.A single softmax attention operation is then applied over all keys and values, enabling pertoken and per-head context-dependent weighting without requiring additional fusion parameters.The inter-layer behavior is controlled through a single hyperparameter, the sliding window size, which allows smooth adjustment between purely linear and full attention while keeping all layers structurally uniform.Experimental results show that NHA surpasses Transformers and other hybrid baselines on recall-intensive and commonsense reasoning tasks.Furthermore, pretrained LLMs can be structurally hybridized with NHA, achieving competitive accuracy while delivering significant efficiency gains.Code is available at https://github.com/JusenD/NHA. Jusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun, Yu Cheng 0001 |
ACL (1) | 4 |
| 2026 | Nirvana: A Specialized Generalist Model With Task-Aware Memory MechanismabstractYuhua Jiang, Shuang Cheng, Yihao Liu, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng, Feifei Gao, Biqing Qi, Bowen Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuhua Jiang, Shuang Cheng, Yihao Liu 0008, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng 0001, Biqing Qi, Bowen Zhou 0002 |
ACL (1) | 6 |
| 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersabstractLinear sequence modeling methods, such as linear attention, state space modeling, and linear RNNs, have recently been recognized as potential alternatives to softmax attention thanks to their linear complexity and competitive performance. However, although their linear-memory advantage during training enables dealing with long sequences, it is still hard to handle extremely long sequences with very limited computational resources. In this paper, we propose Sequence Accumulation (SA) which leverages the common recurrence feature of linear sequence modeling methods to manage infinite context length even on a single GPU. Specifically, SA divides long input sequences into fixed-length sub-sequences and accumulates intermediate states sequentially, which reaches only constant-memory consumption. Additionally, we further propose Sequence Accumulation with Pipeline Parallelism (SAPP), to train large models with infinite context length, without incurring any additional synchronization costs in the sequence dimension. Extensive experiments with a wide range of context lengths are conducted to validate the effectiveness of SA and SAPP on both single and multiple GPUs. Results show that SA and SAPP enable the training of infinite context length on even very limited resources, and are well compatible with the out-of-the-box distributed training techniques. Weigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu Mo |
AAAI | 1 |
| 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language ModelsabstractFaithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on “factual statements” that rephrase source materials while overlooking “cognitive statements” that involve making inferences from the given context. Consequently, evaluating and detecting the hallucination of cognitive statements remains challenging. Inspired by how evidence is assessed in the legal domain, we design a rigorous framework to assess different levels of faithfulness of cognitive statements and introduce the CogniBench dataset where we reveal insightful statistics. To keep pace with rapidly evolving LLMs, we further develop an automatic annotation pipeline that scales easily across different models. This results in a large-scale CogniBench-L dataset, which facilitates training accurate detectors for both factual and cognitive hallucinations. We release our model and datasets at: https://github.com/FUTUREEEEEE/CogniBench Xiaqiang Tang, Keyu Hu, Xi Zhang 0008, Weigao Sun, Sihong Xie |
ACL (1) | 7 |
| 2025 | Liger: Linearizing Large Language Models to Gated Recurrent StructuresabstractTransformers with linear recurrent modeling offer linear-time training and constant-memory inference. Despite their demonstrated efficiency and performance, pretraining such non-standard architectures from scratch remains costly and risky. The linearization of large language models (LLMs) transforms pretrained standard models into linear recurrent structures, enabling more efficient deployment. However, current linearization methods typically introduce additional feature map modules that require extensive fine-tuning and overlook the gating mechanisms used in state-of-the-art linear recurrent models. To address these issues, this paper presents Liger, short for Linearizing LLMs to gated recurrent structures. Liger is a novel approach for converting pretrained LLMs into gated linear recurrent models without adding extra parameters. It repurposes the pretrained key matrix weights to construct diverse gating mechanisms, facilitating the formation of various gated recurrent structures while avoiding the need to train additional components from scratch. Using lightweight fine-tuning with Low-Rank Adaptation (LoRA), Liger restores the performance of the linearized gated recurrent models to match that of the original LLMs. Additionally, we introduce Liger Attention, an intra-layer hybrid attention mechanism, which significantly recovers 93% of the Transformer-based LLM performance at 0.02% pre-training tokens during the linearization process, achieving competitive results across multiple benchmarks, as validated on models ranging from 1B to 8B parameters. Disen Lan, Weigao Sun, Jiaxi Hu, Jusen Du, Yu Cheng 0001 |
ICML | 2 |
| 2025 | Improving Bilinear RNN with Closed-loop ControlabstractRecent efficient sequence modeling methods, such as Gated DeltaNet, TTT, and RWKV-7, have achieved performance improvements by supervising the recurrent memory management through the Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, resulting in a bilinear recursive structure. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then based on the closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on a large-scale corpus. Comba demonstrates its superior performance and computation efficiency on both language modeling and vision tasks. Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang 0002, Weigao Sun |
NeurIPS | 8 |
| 2025 | Multimodal Multi-Agent Joint Traffic Simulation With Collision and Run-Off-Road MitigationabstractSimulation environments featuring data-driven traffic agents have become essential tools for training and assessing autonomous driving (AD) systems, offering a safer and more cost-effective alternative to real-world testing. However, prevalent data-driven simulations often focus on modeling individual traffic participants or indirectly representing their joint future motion distribution in a latent space. These approaches may encounter challenges due to the combinatorial explosion in the number of agents or may lack explainability. To tackle these issues, we introduce a novel learning-based traffic simulation framework aimed at directly modeling joint behaviors in the 2D plane. Our framework utilizes a graph-based scene representation to accommodate an arbitrary number of agents and road elements. It directly generates multiple joint goal sets for all agents in a given scenario and subsequently produces reference lines for each goal set. To enhance interaction with the ego, we design a predictive step simulation module to forecast the ego’s short-term motion and prevent collisions between the step rollout and the prediction during reference line tracking. Experimental results on the large-scale Waymo Open Motion Dataset validate that the proposed model can directly generate multiple socially consistent goal sets, showcasing enhancements in collision avoidance, road compliance, and overall realism. This approach shows promise in advancing the development and deployment of AD technology in a safer and more efficient manner. Xiaoyu Mo, Jintian Ge, Weigao Sun, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Traffic Scene Representation and Encoding With Graph Structure Learning and ExplorationabstractEffective scene representation is critical for trajectory generation tasks in autonomous driving. Existing attention-based methods often rely on fixed-length input elements, limiting their ability to adapt to dynamic traffic environments, and frequently overlook lane connectivity. Methods that consider lane connectivity often segment lanes into smaller pieces, which consequently creates a more complex graph structure with an increased number of nodes and edges. Additionally, many current approaches select agent interactions based on fixed distance thresholds, which may miss important long-range or indirect interactions and ignore the fact that proximity does not indicate interaction. In this paper, we propose SceneGNN, a novel framework for learning traffic scenario representation through interaction graph learning and lane graph exploration. SceneGNN constructs a single heterogeneous graph that integrates both agents and lanes, leveraging high-definition maps without over-segmentation. To model lane connectivity more effectively, we introduce LaneGNN, which combines an Omnidirectional Lane Aggregator (OLA) and Directional Lane Explorer (DLE) to explore lane-to-lane interdependencies. Additionally, instead of relying on proximity-based heuristics, our interaction graph learner dynamically constructs inter-agent edges based on learned features, allowing the model to capture meaningful interactions beyond mere distance. We evaluate SceneGNN on the Waymo Open Motion Dataset, where it achieves competitive performance. Our results demonstrate that by efficiently capturing both agent-lane relationships and inter-agent interactions, SceneGNN improves the accuracy and scalability of multi-agent trajectory prediction for real-world autonomous driving applications. Xiaoyu Mo, Baichuan Lou, Zhiqi Mao, Weigao Sun, Yafei Wang 0001, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Scaling Laws for Linear Complexity Language ModelsabstractThe interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain.In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability.Specifically, we examine the scaling behaviors of three efficient linear architectures.These include TNL (Qin et al., 2024c), a linear attention model with data-independent decay; HGRN2 (Qin et al., 2024e), a linear RNN with data-dependent decay; and cosFormer2 (Qin et al., 2022b(Qin et al., , 2024a)), a linear attention model without decay.We also include LLaMA as a baseline architecture for comparison with softmax attention.These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks.These tasks include validation loss, commonsense reasoning, and information retrieval and generation.The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention. Xuyang Shen, Dong Li 0033, Ruitao Leng, Zhen Qin 0003, Weigao Sun, Yiran Zhong |
EMNLP | 5 |
| 2024 | CO2: Efficient Distributed Training with Full Communication-Computation OverlapabstractThe fundamental success of large language models hinges upon the efficacious implementation of large-scale distributed training techniques. Nevertheless, building a vast, high-performance cluster featuring high-speed communication interconnectivity is prohibitively costly, and accessible only to prominent entities. In this work, we aim to lower this barrier and democratize large-scale training with limited bandwidth clusters. We propose a new approach called CO2 that introduces local-updating and asynchronous communication to the distributed data-parallel training, thereby facilitating the full overlap of COmmunication with COmputation. CO2 is able to attain a high scalability even on extensive multi-node clusters constrained by very limited communication bandwidth. We further propose the staleness gap penalty and outer momentum clipping techniques together with CO2 to bolster its convergence and training stability. Besides, CO2 exhibits seamless integration with well-established ZeRO-series optimizers which mitigate memory consumption of model states with large model training. We also provide a mathematical proof of convergence, accompanied by the establishment of a stringent upper bound. Furthermore, we validate our findings through an extensive set of practical experiments encompassing a wide range of tasks in the fields of computer vision and natural language processing. These experiments serve to demonstrate the capabilities of CO2 in terms of convergence, generalization, and scalability when deployed across configurations comprising up to 128 A100 GPUs. The outcomes emphasize the outstanding capacity of CO2 to hugely improve scalability, no matter on clusters with 800Gbps RDMA or 80Gbps TCP/IP inter-node connections. Weigao Sun, Zhen Qin 0003, Weixuan Sun, Shidi Li, Dong Li 0033, Xuyang Shen, Yu Qiao 0001, Yiran Zhong |
ICLR | 1 |
| 2024 | Various Lengths, Constant Speed: Efficient Language Modeling with Lightning AttentionabstractWe present Lightning Attention, the first linear attention implementation that maintains a constant training speed for various sequence lengths under fixed memory consumption. Due to the issue with cumulative summation operations (cumsum), previous linear attention implementations cannot achieve their theoretical advantage in a casual setting. However, this issue can be effectively solved by utilizing different attention calculation strategies to compute the different parts of attention. Specifically, we split the attention calculation into intra-blocks and inter-blocks and use conventional attention computation for intra-blocks and linear attention kernel tricks for inter-blocks. This eliminates the need for cumsum in the linear attention calculation. Furthermore, a tiling technique is adopted through both forward and backward procedures to take full advantage of the GPU hardware. To enhance accuracy while preserving efficacy, we introduce TransNormerLLM (TNL), a new architecture that is tailored to our lightning attention. We conduct rigorous testing on standard and self-collected datasets with varying model sizes and sequence lengths. TNL is notably more efficient than other language models. In addition, benchmark results indicate that TNL performs on par with state-of-the-art LLMs utilizing conventional transformer structures. The source code is released at github.com/OpenNLPLab/TransnormerLLM. Zhen Qin 0003, Weigao Sun, Dong Li 0033, Xuyang Shen, Weixuan Sun, Yiran Zhong |
ICML | 2 |
| 2020 | pbSGD: Powered Stochastic Gradient Descent Methods for Accelerated Non-Convex OptimizationabstractWe propose a novel technique for improving the stochastic gradient descent (SGD) method to train deep networks, which we term pbSGD. The proposed pbSGD method simply raises the stochastic gradient to a certain power elementwise during iterations and introduces only one additional parameter, namely, the power exponent (when it equals to 1, pbSGD reduces to SGD). We further propose pbSGD with momentum, which we term pbSGDM. The main results of this paper present comprehensive experiments on popular deep learning models and benchmark datasets. Empirical results show that the proposed pbSGD and pbSGDM obtain faster initial training speed than adaptive gradient methods, comparable generalization ability with SGD, and improved robustness to hyper-parameter selection and vanishing gradients. pbSGD is essentially a gradient modifier via a nonlinear transformation. As such, it is orthogonal and complementary to other techniques for accelerating gradient-based optimization such as learning rate schedules. Finally, we show convergence rate analysis for both pbSGD and pbSGDM methods. The theoretical rates of convergence match the best known theoretical rates of convergence for SGD and SGDM methods on nonconvex functions. Beitong Zhou, Jun Liu 0015, Weigao Sun, Ruijuan Chen, Claire J. Tomlin, Ye Yuan 0002 |
IJCAI | 3 |
| 2020 | A Fast Optimal Power Flow Algorithm Using Powerball MethodabstractThe complexity and randomness of the power system with distributed energy resources have led to the difficulties for fast optimal power flow (OPF) analysis. As a remedy, in this paper, we develop an interior point Powerball algorithm to accelerate the OPF solution process. To achieve better convergence characteristics, the proposed IPPB algorithm which is based on the Powerball optimization method, improves the search directions during iterative optimization by a nonlinear transformation. Also, a Newton—Raphson Powerball algorithm is derived for a faster power flow calculation, which is a basic yet critical part of the OPF problem. Numerical case studies are conducted on benchmark power systems with different scales to validate the proposed algorithms. Performances of the proposed algorithms to address improper initial points are studied by randomly picking the initial bus voltages. Numerical study results verify the feasibility and superiority of the proposed algorithms. Hai-Tao Zhang, Weigao Sun, Yuan Zheng Li, Dongfei Fu, Ye Yuan 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Probabilistic Optimal Power Flow With Correlated Wind Power Uncertainty via Markov Chain Quasi-Monte-Carlo SamplingabstractThe irregular and truncated probabilistic characteristics of wind power uncertainty lead to unknown influences on the power system operation. In this article, we propose a new probabilistic optimal power flow (POPF) framework, which can cope with such uncertainties, while taking into account the correlations among the wind generation power in multiple wind farms. A truncated multivariate Gaussian mixture model (Trun-MultiGMM) is designed to describe the irregular and multimodal wind power distributions with its typical truncation feature. Then an efficient Markov chain quasi-Monte-Carlo (MCQMC) sampler is developed to deliver wind power samples from the customized Trun-MultiGMM. Numerical simulations are conducted on the publicly available wind generation datasets and multiple benchmark power systems. The results have verified the effectiveness and efficiency of Trun-MultiGMM as well as the proposed POPF framework with MCQMC sampler. Weigao Sun, Mohsen Zamani, Hai-Tao Zhang, Yuan Zheng Li |
IEEE Trans. Ind. Informatics | 1 |