EDBT 2026 Demo / reviewers in the wild / expert
Simeng Sun
dblp:204/8261
· DBLP profile ↗
20ranked-venue papers
10as first author
16since 2021 · last 2025
0000-0001-9695-9336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingabstractKrishna C Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh, Shantanu Acharya, Somshubra Majumdar, Fei Jia, Samuel Kriman, Simeng Sun, Dima Rekesh, Boris Ginsburg. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Krishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh, Shantanu Acharya, Somshubra Majumdar, Fei Jia, Samuel Kriman, Simeng Sun, Dima Rekesh, Boris Ginsburg |
EMNLP | 9 |
| 2025 | nGPT: Normalized Transformer with Representation Learning on the HypersphereabstractWe propose a novel neural network architecture, the normalized Transformer (nGPT) with representation learning on the hypersphere. In nGPT, all vectors forming the embeddings, MLP, attention matrices and hidden states are unit norm normalized. The input stream of tokens travels on the surface of a hypersphere, with each layer contributing a displacement towards the target output predictions. These displacements are defined by the MLP and attention blocks, whose vector components also reside on the same hypersphere. Experiments show that nGPT learns much faster, reducing the number of training steps required to achieve the same accuracy by a factor of 4 to 20, depending on the sequence length. Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris Ginsburg |
ICLR | 3 |
| 2024 | PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long DocumentsabstractSimeng Sun, Yang Liu, Shuohang Wang, Dan Iter, Chenguang Zhu, Mohit Iyyer. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Simeng Sun, Yang Liu 0124, Shuohang Wang, Dan Iter, Chenguang Zhu 0001, Mohit Iyyer |
EACL (1) | 1 |
| 2024 | TopicGPT: A Prompt-based Topic Modeling FrameworkabstractChau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Alexander Miserlis Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer |
NAACL-HLT | 3 |
| 2023 | Efficiently Upgrading Multilingual Machine Translation Models to Support More LanguagesabstractWith multilingual machine translation (MMT) models continuing to grow in size and number of supported languages, it is natural to reuse and upgrade existing models to save computation as data becomes available in more languages.However, adding new languages requires updating the vocabulary, which complicates the reuse of embeddings.The question of how to reuse existing models while also making architectural changes to provide capacity for both old and new languages has also not been closely studied.In this work, we introduce three techniques that help speed up effective learning of the new languages and alleviate catastrophic forgetting despite vocabulary and architecture mismatches.Our results show that by (1) carefully initializing the network, (2) applying learning rate scaling, and (3) performing data up-sampling, it is possible to exceed the performance of a same-sized baseline model with 30% computation and recover the performance of a larger model trained from scratch with over 50% reduction in computation.Furthermore, our analysis reveals that the introduced techniques help learn the new directions more effectively and alleviate catastrophic forgetting at the same time.We hope our work will guide research into more efficient approaches to growing languages for these MMT models and ultimately maximize the reuse of existing models.* Work done during an internship at Meta AI. Simeng Sun, Maha Elbayad, Anna Y. Sun, James Cross 0003 |
EACL | 1 |
| 2023 | Semantical video coding: Instill static-dynamic clues into structured bitstream for AI tasks
Xin Jin 0014, Ruoyu Feng 0001, Simeng Sun, Runsen Feng, Tianyu He, Zhibo Chen 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | GraphIQA: Learning Distortion Graph Representations for Blind Image Quality AssessmentabstractA good distortion representation is crucial for the success of deep blind image quality assessment (BIQA). However, most previous methods do not effectively model the relationship between distortions or the distribution of samples with the same distortion type but different distortion levels. In this work, we start from the analysis of the relationship between perceptual image quality and distortion-related factors, such as distortion types and levels. Then, we propose a Distortion Graph Representation (DGR) learning framework for IQA, named GraphIQA, in which each distortion is represented as a graph,i.e., DGR. One can distinguish distortion types by learning the contrast relationship between these different DGRs, and can infer the ranking distribution of samples from different levels in a DGR. Specifically, we develop two sub-networks to learn the DGRs: a) Type Discrimination Network (TDN) that aims to embed DGR into a compact code for better discriminating distortion types and learning the relationship between types; b) Fuzzy Prediction Network (FPN) that aims to extract the distributional characteristics of the samples in a DGR and predicts fuzzy degrees based on a Gaussian prior. Experiments show that our GraphIQA achieves state-of-the-art performance on many benchmark datasets of both synthetic and authentic distortions. Simeng Sun, Tao Yu 0012, Jiahua Xu 0001, Wei Zhou 0021, Zhibo Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Alternative Input Signals Ease Transfer in Multilingual Machine TranslationabstractSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán |
ACL (1) | 1 |
| 2022 | Image Coding for Machines with Omnipotent Feature Learning
Ruoyu Feng 0001, Xin Jin 0014, Zongyu Guo, Runsen Feng, Tianyu He, Zhizheng Zhang 0004, Simeng Sun, Zhibo Chen 0001 |
ECCV (37) | 8 |
| 2022 | ChapterBreak: A Challenge Dataset for Long-Range Language ModelsabstractWhile numerous architectures for long-range language models (LRLMs) have recently been proposed, a meaningful evaluation of their discourse-level language understanding capabilities has not yet followed.To this end, we introduce CHAPTERBREAK, a challenge dataset that provides an LRLM with a long segment from a narrative that ends at a chapter boundary and asks it to distinguish the beginning of the ground-truth next chapter from a set of negative segments sampled from the same narrative.A fine-grained human annotation reveals that our dataset contains many complex types of chapter transitions (e.g., parallel narratives, cliffhanger endings) that require processing global context to comprehend.Experiments on CHAPTERBREAK show that existing LRLMs fail to effectively leverage long-range context, substantially underperforming a segment-level model trained directly for this task.We publicly release our CHAPTERBREAK dataset to spur more principled future research into LRLMs. 1 Simeng Sun, Katherine Thai, Mohit Iyyer |
NAACL-HLT | 1 |
| 2021 | Learning Omni-Frequency Region-adaptive Representations for Real Image Super-ResolutionabstractTraditional single image super-resolution (SISR) methods that focus on solving single and uniform degradation (i.e., bicubic down-sampling), typically suffer from poor performance when applied into real-world low-resolution (LR) images due to the complicated realistic degradations. The key to solving this more challenging real image super-resolution (RealSR) problem lies in learning feature representations that are both informative and content-aware. In this paper, we propose a Omni-frequency Region-adaptive Network (OR-Net) to address both challenges, here we call features of all low, middle and high frequencies omni-frequency features. Specifically, we start from the frequency perspective and design a Frequency Decomposition (FD) module to separate different frequency components to comprehensively compensate the information lost for real LR image. Then, considering the different regions of real LR image have different frequency information lost, we further design a Region-adaptive Frequency Aggregation (RFA) module by leveraging dynamic convolution and spatial attention to adaptively restore frequency components for different regions. The extensive experiments endorse the high-efficient, effective, and scenario-agnostic nature of our OR-Net for RealSR. Xin Li 0082, Xin Jin 0014, Tao Yu 0012, Simeng Sun, Yingxue Pang, Zhizheng Zhang 0004, Zhibo Chen 0001 |
AAAI | 4 |
| 2021 | Energy-Based Reranking: Improving Neural Machine Translation Using Energy-Based ModelsabstractSumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun, Mohit Iyyer, Andrew McCallum. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun, Mohit Iyyer, Andrew McCallum |
ACL/IJCNLP (1) | 4 |
| 2021 | Do Long-Range Language Models Actually Use Long-Range Context?abstractLanguage models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions.Recent efforts to improve the efficiency of self-attention have led to a proliferation of long-range Transformer language models, which can process much longer sequences than models of the past.However, the ways in which such models take advantage of the longrange context remain unclear.In this paper, we perform a fine-grained analysis of two longrange Transformer language models (including the Routing Transformer, which achieves state-of-the-art perplexity on the PG-19 longsequence LM benchmark dataset) that accept input sequences of up to 8K tokens.Our results reveal that providing long-range context (i.e., beyond the previous 2K tokens) to these models only improves their predictions on a small set of tokens (e.g., those that can be copied from the distant context) and does not help at all for sentence-level prediction tasks.Finally, we discover that PG-19 contains a variety of different document types and domains, and that long-range context helps most for literary novels (as opposed to textbooks or magazines). Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit Iyyer |
EMNLP (1) | 1 |
| 2021 | IGA: An Intent-Guided Authoring AssistantabstractSimeng Sun, Wenlong Zhao, Varun Manjunatha, Rajiv Jain, Vlad Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Simeng Sun, Wenlong Zhao 0001, Varun Manjunatha, Rajiv Jain, Vlad I. Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer |
EMNLP (1) | 1 |
| 2021 | Revisiting Simple Neural Probabilistic Language ModelsabstractRecent progress in language modeling has been driven not only by advances in neural architectures, but also through hardware and optimization improvements.In this paper, we revisit the neural probabilistic language model (NPLM) of Bengio et al. (2003), which simply concatenates word embeddings within a fixed window and passes the result through a feed-forward network to predict the next word.When scaled up to modern hardware, this model (despite its many limitations) performs much better than expected on word-level language model benchmarks.Our analysis reveals that the NPLM achieves lower perplexity than a baseline Transformer with short input contexts but struggles to handle long-term dependencies.Inspired by this result, we modify the Transformer by replacing its first selfattention layer with the NPLM's local concatenation layer, which results in small but consistent perplexity decreases across three wordlevel language modeling datasets. Simeng Sun, Mohit Iyyer |
NAACL-HLT | 1 |
| 2021 | Semantic Structured Image Coding Framework for Multiple Intelligent ApplicationsabstractFast-growing intelligent media processing applications demand efficient processing throughout the processing chain from the edge to the cloud, and the complexity bottleneck usually lies in the parallel decoding of multiple-channel compressed bitstreams before analyzing. This occurs because the traditional media coding scheme generates a binary stream without a semantic structure, which is unable to be operated directly at the bitstream level to support different tasks such as classification, recognition, detection, etc. Therefore, in this article, we propose a learning-based semantically structured image coding (SSIC) framework to generate a semantically structured bitstream (SSB), where each part of the bitstream represents a specific object and can be directly used for the aforementioned intelligent tasks. Specifically, we integrate an object location extraction module into the compression framework to locate and align objects in the feature domain. Then, each object together with the background is compressed separately and reorganized to form a structured bitstream to enable the analysis or reconstruction of specific objects directly from partial bitstream. Furthermore, in contrast to existing learning-based compression schemes that train the specific model for a specific bitrate, we share most of the model parameters among various bitrates to significantly reduce the model size for variable-rate compression. The experimental results demonstrate the effectiveness of the proposed coding scheme whose compression performance is comparable to existing image coding schemes, where intelligent tasks such as classification and pose estimation can be directly performed on a partial bitstream without performance degradation, significantly reducing the complexity for analyzing tasks. Simeng Sun, Tianyu He, Zhibo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Hard-Coded Gaussian Attention for Neural Machine TranslationabstractRecent work has questioned the importance of the Transformer's multi-headed attention for achieving high translation quality.We push further in this direction by developing a "hardcoded" attention variant without any learned parameters.Surprisingly, replacing all learned self-attention heads in the encoder and decoder with fixed, input-agnostic Gaussian distributions minimally impacts BLEU scores across four different language pairs.However, additionally hard-coding cross attention (which connects the decoder to the encoder) significantly lowers BLEU, suggesting that it is more important than self-attention.Much of this BLEU drop can be recovered by adding just a single learned cross attention head to an otherwise hard-coded Transformer.Taken as a whole, our results offer insight into which components of the Transformer are actually important, which we hope will guide future work into the development of simpler and more efficient attention-based models. Weiqiu You, Simeng Sun, Mohit Iyyer |
ACL | 2 |
| 2019 | The Feasibility of Embedding Based Automatic Evaluation for Single Document SummarizationabstractSimeng Sun, Ani Nenkova. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Simeng Sun, Ani Nenkova |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Beyond Coding: Detection-driven Image Compression with Semantically Structured Bit-streamabstractWith the development of 5G and edge computing, it is increasingly important to offload intelligent media computing to edge device. Traditional media coding scheme codes the media into one binary stream without a semantic structure, which prevents many important intelligent applications from operating directly in bit-stream level, including semantic analysis, parsing specific content, media editing, etc. Therefore, in this paper, we propose a learning based Semantically Structured Coding (SSC) framework to generate Semantically Structured Bit-stream (SSB), where each part of bit-stream represents a certain object and can be directly used for aforementioned tasks. Specifically, we integrate an object detection module in our compression framework to locate and align the object in feature domain. After applying quantization and entropy coding, the features are re-organized according to detected and aligned objects to form a bit-stream. Besides, different from existing learning-based compression schemes that individually train models for specific bit-rate, we share most of model parameters among various bit-rates to significantly reduce model size for variable-rate compression. Experimental results demonstrate that only at the cost of negligible overhead, objects can be completely reconstructed from partial bit-stream. We also verified that classification and pose estimation can be directly performed on partial bit-stream without performance degradation. Tianyu He, Simeng Sun, Zongyu Guo, Zhibo Chen 0001 |
PCS | 2 |
| 2019 | System architecture for high-performance permissioned blockchains
Libo Feng, Hui Zhang 0028, Wei-Tek Tsai, Simeng Sun |
Frontiers Comput. Sci. | 4 |