VLDB 2026 Research / reviewers in the wild / expert
Yikang Shen
dblp:152/8226
· DBLP profile ↗
47ranked-venue papers
9as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 9 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Unified Expertise: Learning a Single Vision Model From Diverse PerceptionabstractMulti-task learning (MTL) presents greater optimization challenges than single-task learning (STL) due to conflicting gradients across tasks. While parameter sharing promotes cooperation among related tasks, many tasks require specialized representations. To balance cooperation and specialization, we propose Mod-Squad (Chen et al. 2023), a modular transformer-based model composed of a "squad" of experts. Each task activates a sparse subset of experts through a differentiable matching process, guided by a novel mutual information-based loss. This modular structure avoids full backbone sharing and scales effectively with the number of tasks and dataset size. In this extended version, we generalize Mod-Squad to support multi-dataset pre-training, enabling joint learning across disjoint, single-task datasets (e.g., ImageNet, COCO, ADE20 K). This is achieved via a new formulation of the mutual information loss that unifies learning across heterogeneous sources. More importantly, while most prior work in large models has focused on efficiency, few have explored adjustable efficiency. In this study, we further evaluate the model's generalization to downstream tasks and introduce a set of efficient adaptation techniques that leverage Mod-Squad's modularity for flexible fine-tuning-enabling dynamic adjustment of model size, parameter count, and computational cost. Additionally, we present a hybrid adaptation scheme that combines these techniques to achieve favorable performance-efficiency trade-offs. In summary, Mod-Squad provides a robust foundation for sparse modular models that can learn from diverse supervision and datasets. Its emergent modularity enables strong generalization, decomposition into high-performing components, and rapid, resource-efficient adaptation for downstream applications. Zitian Chen, Mingyu Ding, Yikang Shen, Erik G. Learned-Miller, Chuang Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | LaMAGIC: Advanced Circuit Formulations for Language-Model-based Topology Generation for Analog Integrated CircuitsabstractIn the realm of electronic and electrical engineering, automation of analog circuit is increasingly vital given the complexity and customized requirements of modern applications. However, existing methods only develop search-based algorithms that require many simulation iterations to design a custom circuit topology, which is usually a time-consuming process. To this end, we introduce LaMAGIC, a language model-based topology generation model that leverages supervised finetuning for automated analog circuit design. LaMAGIC can efficiently generate an optimized circuit design from the custom specification in a single pass. The generated circuit is validated by the simulator to meet the performance requirement with high precision. Our approach involves a meticulous development and analysis of various input and output formulations for circuit. These formulations can ensure canonical representations and align with the autoregressive nature of LMs for representing analog circuits as graphs. In addition, our novel transformer model supports float-input to effectively learn the mapping between numerical performance and circuits. The experimental results show that LaMAGIC achieves a success rate of up to 96% under a strict tolerance of 0.01. Also, we examine the scalability and adaptability of LaMAGIC under scarce data scenario on more complex circuits. Our findings reveal the enhanced effectiveness of our succinct float-input canonical formulation with identifier, suggesting its suitability for handling intricate circuits. Our ablation study evaluates various design choices of LM training and inference, providing insights for future domain-specific generation tasks. This research not only demonstrates the potential of language models in graph generation, but also builds a foundational framework for future explorations in automated analog circuit design. Chen-Chia Chang, Wan-Hsuan Lin, Yikang Shen, Guanglei Zhou, Yiran Chen 0001, Xin Zhang 0025 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | API Pack: A Massive Multi-Programming Language Dataset for API Call GenerationabstractWe introduce API Pack, a massive multi-programming language dataset containing over one million instruction-API calls for improving the API call generation capabilities of large language models. Our evaluation highlights three key findings: First, fine-tuning on API Pack enables open-source models to outperform GPT-3.5 and GPT-4 in generating code for entirely new API calls. We show this by fine-tuning CodeLlama-13B on 20,000 Python instances from API Pack. Second, fine-tuning on a large dataset in one language, combined with smaller datasets from others, improves API generation accuracy across multiple languages. Third, we confirm the benefits of larger datasets for API generalization, as increasing fine-tuning data to one million instances enhances generalization to new APIs. To support further research, we open-source the API Pack dataset, trained model, and code at https://github.com/zguo0525/API-Pack. Adriana Meza Soria, Yikang Shen, Rameswar Panda |
ICLR | 4 |
| 2025 | Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth StudyabstractThe self-attention mechanism traditionally relies on the softmax operator, necessitating positional embeddings like RoPE, or position biases to account for token order.
But current methods using still face length generalisation challenges.
We investigate an alternative attention mechanism based on the stick-breaking process in larger scale settings.
The method works as follows: For each token before the current, we determine a break point, which represents the proportion of the stick, the weight of the attention, to allocate to the current token.
We repeat this on the remaining stick, until all tokens are allocated a weight, resulting in a sequence of attention weights.
This process naturally incorporates recency bias, which has linguistic motivations for grammar parsing (Shen et al., 2017).
We study the implications of replacing the conventional softmax-based attention mechanism with stick-breaking attention.
We then discuss implementation of numerically stable stick-breaking attention and adapt Flash Attention to accommodate this mechanism.
When used as a drop-in replacement for current softmax+RoPE attention systems, we find that stick-breaking attention performs competitively with current methods on length generalisation and downstream tasks.
Stick-breaking also performs well at length generalisation, allowing a model trained with $2^{11}$ context window to perform well at $2^{14}$ with perplexity improvements. Shawn Tan, Aaron C. Courville, Rameswar Panda, Yikang Shen |
ICLR | 5 |
| 2025 | LaMAGIC2: Advanced Circuit Formulations for Language Model-Based Analog Topology GenerationabstractAutomation of analog topology design is crucial due to customized requirements of modern applications with heavily manual engineering efforts.
The state-of-the-art work applies a sequence-to-sequence approach and supervised finetuning on language models to generate topologies given user specifications.
However, its circuit formulation is inefficient due to $O(|V|^2)$ token length and suffers from low precision sensitivity to numeric inputs.
In this work, we introduce LaMAGIC2, a succinct float-input canonical formulation
with identifier (SFCI) for language model-based analog topology generation.
SFCI addresses these challenges by improving component-type recognition through identifier-based representations, reducing token length complexity to $O(|V|)$, and enhancing numeric precision sensitivity for better performance under tight tolerances.
Our experiments demonstrate that LaMAGIC2 achieves 34\% higher success rates under a tight tolerance 0.01 and 10X lower MSEs compared to a prior method.
LaMAGIC2 also exhibits better transferability for circuits with more vertices with up to 58.5\% improvement.
These advancements establish LaMAGIC2 as a robust framework for analog topology generation. Chen-Chia Chang, Wan-Hsuan Lin, Yikang Shen, Yiran Chen 0001, Xin Zhang 0025 |
ICML | 3 |
| 2025 | PaTH Attention: Position Encoding via Accumulating Householder TransformationsabstractThe attention mechanism is a core primitive in modern large language models (LLMs) and AI more broadly. Since attention by itself is permutation-invariant, position encoding is essential for modeling structured domains such as language. Rotary position encoding (RoPE) has emerged as the de facto standard approach for position encoding and is part of many modern LLMs. However, in RoPE the key/query transformation between two elements in a sequence is only a function of their relative position and otherwise independent of the actual input. This limits the expressivity of RoPE-based transformers.
This paper describes PaTH, a flexible data-dependent position encoding scheme based on accumulated products of Householder(like) transformations, where each transformation is data-dependent, i.e., a function of the input. We derive an efficient parallel algorithm for training through exploiting a compact representation of products of Householder matrices, and implement a FlashAttention-style blockwise algorithm. Across both targeted synthetic benchmarks and moderate-scale real-world language modeling experiments, we find that PaTH improves upon RoPE and other recent baselines. Finally, we show that we can convert pretrained RoPE transformers into PaTH with continued pretraining. Yikang Shen, Kaiyue Wen, Shawn Tan, Liliang Ren, Rameswar Panda |
NeurIPS | 2 |
| 2024 | Visual Chain-of-Thought Prompting for Knowledge-Based Visual ReasoningabstractKnowledge-based visual reasoning remains a daunting task since it not only requires machines to interpret the concepts and relationships from visual scenes but also associate them with external world knowledge to conduct a chain of reasoning on open-world questions. Previous works, however, treat visual perception and language-based reasoning as two independent modules, failing to attend to both modules throughout all stages of reasoning. To this end, we propose Visual Chain-of-thought Prompting (VCTP) for knowledge-based reasoning, which involves the interaction between visual content and natural language in an iterative step-by-step reasoning manner. VCTP contains three stages, see, think, and confirm. The see stage scans the image and grounds the visual concept candidates with a visual perception model. The think stage adopts a pre-trained large language model (LLM) to attend to key visual concepts from natural language questions adaptively. It then transforms key visual context into text context for prompting with a visual captioning model, and adopts the LLM to generate the answer. The confirm stage further uses the LLM to generate the supporting rationale to the answer, which is then passed through a cross-modality classifier to verify that it’s consistent with the visual context. We iterate through the think-confirm stages to ensure the verified rationale is consistent with the answer. We conduct experiments on a range of knowledge-based visual reasoning datasets. We found our VCTP enjoys several benefits, 1). it achieves better performance than the previous few-shot learning baselines; 2). it enjoys the total transparency and trustworthiness of the whole reasoning process by providing rationales for each reasoning step; 3). it is computation-efficient compared with other fine-tuning baselines. Our code is available at https://github.com/UMass-Foundation-Model/VisualCoT.git Zhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong, Zhiqing Sun, Dan Gutfreund, Chuang Gan 0001 |
AAAI | 3 |
| 2024 | FlexAttention for Efficient High-Resolution Vision-Language Models
Delin Chen, Tianle Cai, Peihao Chen, Yining Hong, Zhenfang Chen, Yikang Shen, Chuang Gan 0001 |
ECCV (25) | 7 |
| 2024 | The Consensus Game: Language Model Generation via Equilibrium SearchabstractWhen applied to question answering and other text generation tasks, language models (LMs) may be queried generatively (by sampling answers from their output distribution) or discriminatively (by using them to score or rank a set of candidate answers). These procedures sometimes yield very different predictions. How do we reconcile mutually incompatible scoring procedures to obtain coherent LM predictions? We introduce a new, a training-free, game-theoretic procedure for language model decoding. Our approach casts language model decoding as a regularized imperfect-information sequential signaling game—which we term the concensus game—in which a generator seeks to communicate an abstract correctness parameter using natural language sentences to a discriminator. We develop computational procedures for finding approximate equilibria of this game, resulting in a decoding algorithm we call equilibrium-ranking. Applied to a large number of tasks (including reading comprehension, commonsense reasoning, mathematical problem-solving, and assistive dialog), equilibrium-ranking consistently improves performance over existing LM decoding procedures. These improvements are sometimes substantial—on multiple benchmarks, we observe that applying equilibrium-ranking to LLaMA-7B outperforms the much larger LLaMA-65B and PaLM-540B models. Athul Paul Jacob, Yikang Shen, Gabriele Farina, Jacob Andreas |
ICLR | 2 |
| 2024 | CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative DecodingabstractA remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large vision-language foundation models (VLMs) fall short of such compositional abilities due to their ``bag-of-words" behaviors and inability to construct words that correctly represent visual entities and the relations among the entities. To this end, we propose CoVLM, which can guide the LLM to explicitly compose visual entities and relationships among the text and dynamically communicate with the vision encoder and detection network to achieve vision-language communicative decoding. Specifically, we first devise a set of novel communication tokens for the LLM, for dynamic communication between the visual detection system and the language system. A communication token is generated by the LLM following a visual entity or a relation, to inform the detection network to propose regions that are relevant to the sentence generated so far. The proposed regions-of-interests (ROIs) are then fed back into the LLM for better language generation contingent on the relevant regions. The LLM is thus able to compose the visual entities and relationships through the communication tokens. The vision-to-language and language-to-vision communication are iteratively performed until the entire sentence is generated. Our framework seamlessly bridges the gap between visual perception and LLMs and outperforms previous VLMs by a large margin on compositional reasoning benchmarks (e.g., ~20% in HICO-DET mAP, ~14% in Cola top-1 accuracy, and ~3% on ARO top-1 accuracy). We also achieve state-of-the-art performances on traditional vision-language tasks such as referring expression comprehension and visual question answering. Delin Chen, Yining Hong, Zhenfang Chen, Peihao Chen, Yikang Shen, Chuang Gan 0001 |
ICLR | 6 |
| 2024 | SALMON: Self-Alignment with Instructable Reward ModelsabstractSupervised Fine-Tuning (SFT) on response demonstrations combined with Reinforcement Learning from Human Feedback (RLHF) constitutes a powerful paradigm for aligning LLM-based AI agents. However, a significant limitation of such an approach is its dependency on high-quality human annotations, making its application to intricate tasks challenging due to difficulties in obtaining consistent response demonstrations and in-distribution response preferences. This paper presents a novel approach, namely SALMON, to align base language models with minimal human supervision, using only a small set of human-defined principles, yet achieving superior performance. Central to our approach is an instructable reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. By merely adjusting these principles during the RL training phase, we gain full control over the preferences with the instructable reward model, subsequently influencing the behavior of the RL-trained policy models, and reducing the reliance on the collection of online human preferences. Applying our method to the LLaMA-2-70b base language model, we developed an AI assistant named Dromedary-2. With only 6 exemplars for in-context learning and 31 human-defined principles, Dromedary-2 significantly surpasses the performance of several state-of-the-art AI systems, including LLaMA-2-Chat-70b, on various benchmark datasets. We have open-sourced the code and model weights to encourage further research into aligning LLM-based AI agents with enhanced supervision efficiency, improved controllability, and scalable oversight. Zhiqing Sun, Yikang Shen, Qinhong Zhou, Zhenfang Chen, David D. Cox, Yiming Yang 0002, Chuang Gan 0001 |
ICLR | 2 |
| 2024 | LaMAGIC: Language-Model-based Topology Generation for Analog Integrated CircuitsabstractIn the realm of electronic and electrical engineering, automation of analog circuit is increasingly vital given the complexity and customized requirements of modern applications. However, existing methods only develop search-based algorithms that require many simulation iterations to design a custom circuit topology, which is usually a time-consuming process. To this end, we introduce LaMAGIC, a pioneering language model-based topology generation model that leverages supervised finetuning for automated analog circuit design. LaMAGIC can efficiently generate an optimized circuit design from the custom specification in a single pass. Our approach involves a meticulous development and analysis of various input and output formulations for circuit. These formulations can ensure canonical representations of circuits and align with the autoregressive nature of LMs to effectively addressing the challenges of representing analog circuits as graphs. The experimental results show that LaMAGIC achieves a success rate of up to 96% under a strict tolerance of 0.01. We also examine the scalability and adaptability of LaMAGIC, specifically testing its performance on more complex circuits. Our findings reveal the enhanced effectiveness of our adjacency matrix-based circuit formulation with floating-point input, suggesting its suitability for handling intricate circuit designs. This research not only demonstrates the potential of language models in graph generation, but also builds a foundational framework for future explorations in automated analog circuit design. Chen-Chia Chang, Yikang Shen, Shaoze Fan, Jing Li 0025, Ningyuan Cao, Yiran Chen 0001, Xin Zhang 0025 |
ICML | 2 |
| 2024 | Gated Linear Attention Transformers with Hardware-Efficient TrainingabstractTransformers with linear attention allow for efficient parallel training but can simultaneously be formulated as an RNN with 2D (matrix-valued) hidden states, thus enjoying linear-time inference complexity. However, linear attention generally underperforms ordinary softmax attention. Moreover, current implementations of linear attention lack I/O-awareness and are thus slower than highly optimized implementations of softmax attention. This work describes a hardware-efficient algorithm for linear attention that trades off memory movement against parallelizability. The resulting implementation, dubbed FlashLinearAttention, is faster than FlashAttention-2 as a standalone layer even on short sequence lengths (e.g., 1K). We then generalize this algorithm to a more expressive variant of linear attention with data-dependent gates. When used as a replacement for the standard attention layer in Transformers, the resulting gated linear attention (GLA) Transformer is found to perform competitively against the LLaMA-architecture Transformer as well recent linear-time-inference baselines such as RetNet and Mamba on moderate-scale language modeling experiments. GLA Transformer is especially effective at length generalization, enabling a model trained on 2K to generalize to sequences longer than 20K without significant perplexity degradations. For training speed, the GLA Transformer has higher throughput than a similarly-sized Mamba model. Bailin Wang, Yikang Shen, Rameswar Panda |
ICML | 3 |
| 2024 | Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingabstractLLMs are computationally expensive to pre-train due to their large scale.
Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones.
However, the viability of these model growth methods in efficient LLM pre-training remains underexplored.
This work identifies three critical $\underline{\textit{O}}$bstacles: ($\textit{O}$1) lack of comprehensive evaluation, ($\textit{O}$2) untested viability for scaling, and ($\textit{O}$3) lack of empirical guidelines.
To tackle $\textit{O}$1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting.
Our findings reveal that a depthwise stacking operator, called $G_{\text{stack}}$, exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines.
Motivated by these promising results, we conduct extensive experiments to delve deeper into $G_{\text{stack}}$ to address $\textit{O}$2 and $\textit{O}$3.
For $\textit{O}$2 (untested scalability), our study shows that $G_{\text{stack}}$ is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens.
For example, compared to a conventionally trained 7B model using 300B tokens, our $G_{\text{stack}}$ model converges to the same loss with 194B tokens, resulting in a 54.6\% speedup.
We further address $\textit{O}$3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for $G_{\text{stack}}$, making it practical in general LLM pre-training.
We also provide in-depth discussions and comprehensive ablation studies of $G_{\text{stack}}$.
Our code and pre-trained model are available at https://llm-stacking.github.io/. Wenyu Du, Tongxu Luo, Zihan Qiu, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu 0001 |
NeurIPS | 5 |
| 2024 | Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionabstractCurrent AI alignment methodologies rely on human-provided demonstrations or judgments, and the learned capabilities of AI systems would be upper-bounded by human capabilities as a result. This raises a challenging research question: How can we keep improving the systems when their capabilities have surpassed the levels of humans? This paper answers this question in the context of tackling hard reasoning tasks (e.g., level 4-5 MATH problems) via learning from human annotations on easier tasks (e.g., level 1-3 MATH problems), which we term as easy-to-hard generalization. Our key insight is that an evaluator (reward model) trained on supervisions for easier tasks can be effectively used for scoring candidate solutions of harder tasks and hence facilitating easy-to-hard generalization over different levels of tasks. Based on this insight, we propose a novel approach to scalable alignment, which firstly trains the (process-supervised) reward models on easy problems (e.g., level 1-3), and then uses them to evaluate the performance of policy models on hard problems. We show that such easy-to-hard generalization from evaluators can enable easy-to-hard generalizations in generators either through re-ranking or reinforcement learning (RL). Notably, our process-supervised 7b RL model and 34b model (reranking@1024) achieves an accuracy of 34.0% and 52.5% on MATH500, respectively, despite only using human supervision on easy problems. Our approach suggests a promising path toward AI systems that advance beyond the frontier of human supervision. Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang 0002, Sean Welleck, Chuang Gan 0001 |
NeurIPS | 3 |
| 2024 | Parallelizing Linear Transformers with the Delta Rule over Sequence LengthabstractTransformers with linear attention (i.e., linear transformers) and state-space models have recently been suggested as a viable linear-time alternative to transformers with softmax attention. However, these models still underperform transformers especially on tasks that require in-context retrieval. While more expressive variants of linear transformers which replace the additive update in linear transformers with the delta rule (DeltaNet) have been found to be more effective at associative recall, existing algorithms for training such models do not parallelize over sequence length and are thus inefficient to train on modern hardware. This work describes a hardware-efficient algorithm for training linear transformers with the delta rule, which exploits a memory-efficient representation for computing products of Householder matrices. This algorithm allows us to scale up DeltaNet to standard language modeling settings. We train a 1.3B model for 100B tokens and find that it outperforms recent linear-time baselines such as Mamba and GLA in terms of perplexity and zero-shot performance on downstream tasks. We also experiment with two hybrid models which combine DeltaNet layers with (1) sliding-window attention layers every other layer or (2) two global attention layers, and find that these hybrids outperform strong transformer baselines. Bailin Wang, Yikang Shen |
NeurIPS | 4 |
| 2023 | Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task LearnersabstractOptimization in multi-task learning (MTL) is more challenging than single-task learning (STL), as the gradient from different tasks can be contradictory. When tasks are related, it can be beneficial to share some parameters among them (cooperation). However, some tasks require additional parameters with expertise in a specific type of data or discrimination (specialization). To address the MTL challenge, we propose Mod-Squad, a new model that is Modularized into groups of experts (a ‘Squad’). This structure allows us to formalize cooperation and specialization as the process of matching experts and tasks. We optimize this matching process during the training of a single model. Specifically, we incorporate mixture of experts (MoE) layers into a transformer model, with a new loss that incorporates the mutual dependence between tasks and experts. As a result, only a small set of experts are activated for each task. This prevents the sharing of the entire backbone model between all tasks, which strengthens the model, especially when the training set size and the number of tasks scale up. More interestingly, for each task, we can extract the small set of experts as a standalone model that maintains the same performance as the large model. Extensive experiments on the Taskonomy dataset with 13 vision tasks and the PASCAL-Context dataset with 5 vision tasks show the superiority of our approach. The project page can be accessed at https://vis-www.cs.umass.edu/Mod-Squad. Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik G. Learned-Miller, Chuang Gan 0001 |
CVPR | 2 |
| 2023 | Visual Dependency Transformers: Dependency Tree Emerges from Reversed AttentionabstractHumans possess a versatile mechanism for extracting structured representations of our visual world. When looking at an image, we can decompose the scene into entities and their parts as well as obtain the dependencies between them. To mimic such capability, we propose Visual Dependency Transformers (DependencyViT)11https://github.com/dingmyu/DependencyViT that can induce visual dependencies without any labels. We achieve that with a novel neural operator called reversed attention that can naturally capture long-range visual dependencies between image patches. Specifically, we formulate it as a dependency graph where a child token in reversed attention is trained to attend to its parent tokens and send information following a normalized probability distribution rather than gathering information in conventional self-attention. With such a design, hierarchies naturally emerge from reversed attention layers, and a dependency tree is progressively induced from leaf nodes to the root node unsupervisedly. DependencyViT offers several appealing benefits. (i) Entities and their parts in an image are represented by different subtrees, enabling part partitioning from dependencies; (ii) Dynamic visual pooling is made possible. The leaf nodes which rarely send messages can be pruned without hindering the model performance, based on which we propose the lightweight DependencyViT-Lite to reduce the computational and memory footprints; (iii) DependencyViT works well on both self- and weakly-supervised pretraining paradigms on ImageNet, and demonstrates its effectiveness on 8 datasets and 5 tasks, such as unsupervised part and saliency segmentation, recognition, and detection. Mingyu Ding, Yikang Shen, Lijie Fan, Zhenfang Chen, Zitian Chen, Ping Luo 0002, Josh Tenenbaum, Chuang Gan 0001 |
CVPR | 2 |
| 2023 | Sparse Universal TransformerabstractThe Universal Transformer (UT) is a variant of the Transformer that shares parameters across its layers.Empirical evidence shows that UTs have better compositional generalization than Vanilla Transformers (VTs) in formal language tasks.The parameter-sharing also affords it better parameter efficiency than VTs.Despite its many advantages, scaling UT parameters is much more compute and memory intensive than scaling up a VT.This paper proposes the Sparse Universal Transformer (SUT), which leverages Sparse Mixture of Experts (SMoE) and a new stick-breaking-based dynamic halting mechanism to reduce UT's computation complexity while retaining its parameter efficiency and generalization ability.Experiments show that SUT achieves the same performance as strong baseline models while only using half computation and parameters on WMT'14 and strong generalization results on formal language tasks (Logical inference and CFQ).The new halting mechanism also enables around 50% reduction in computation during inference with very little performance decrease on formal language tasks. Shawn Tan, Yikang Shen, Zhenfang Chen, Aaron C. Courville, Chuang Gan 0001 |
EMNLP | 2 |
| 2023 | TextPSG: Panoptic Scene Graph Generation from Textual DescriptionsabstractPanoptic Scene Graph has recently been proposed for comprehensive scene understanding. However, previous works adopt a fully-supervised learning manner, requiring large amounts of pixel-wise densely-annotated data, which is always tedious and expensive to obtain. To address this limitation, we study a new problem of Panoptic Scene Graph Generation from Purely Textual Descriptions (Caption-to-PSG). The key idea is to leverage the large collection of free image-caption data on the Web alone to generate panoptic scene graphs. The problem is very challenging for three constraints: 1) no location priors; 2) no explicit links between visual regions and textual entities; and 3) no predefined concept sets. To tackle this problem, we propose a new framework TextPSG consisting of four modules, i.e., a region grouper, an entity grounder, a segment merger, and a label generator, with several novel techniques. The region grouper first groups image pixels into different segments and the entity grounder then aligns visual segments with language entities based on the textual description of the segment being referred to. The grounding results can thus serve as pseudo labels enabling the segment merger to learn the segment similarity as well as guiding the label generator to learn object semantics and relation predicates, resulting in a fine-grained structured scene understanding. Our framework is effective, significantly outperforming the baselines and achieving strong out-of-distribution robustness. We perform comprehensive ablation studies to corroborate the effectiveness of our design choices and provide an in-depth analysis to highlight future directions. Our code, data, and results are available on our project page: https://vis-www.cs.umass.edu/TextPSG. Chengyang Zhao, Yikang Shen, Zhenfang Chen, Mingyu Ding, Chuang Gan 0001 |
ICCV | 2 |
| 2023 | Transformer-Patcher: One Mistake Worth One Neuron
Yikang Shen, Xiaofeng Zhang 0004, Wenge Rong, Zhang Xiong 0001 |
ICLR | 2 |
| 2023 | Hyper-Decision Transformer for Efficient Online Policy Adaptation
Mengdi Xu, Yikang Shen, Ding Zhao, Chuang Gan 0001 |
ICLR | 3 |
| 2023 | Planning with Large Language Models for Code Generation
Zhenfang Chen, Yikang Shen, Mingyu Ding, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 3 |
| 2023 | Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionabstractRecent AI-assistant agents, such as ChatGPT, predominantly rely on supervised fine-tuning (SFT) with human annotations and reinforcement learning from human feedback (RLHF) to align the output of large language models (LLMs) with human intentions, ensuring they are helpful, ethical, and reliable. However, this dependence can significantly constrain the true potential of AI-assistant agents due to the high cost of obtaining human supervision and the related issues on quality, reliability, diversity, self-consistency, and undesirable biases. To address these challenges, we propose a novel approach called SELF-ALIGN, which combines principle-driven reasoning and the generative power of LLMs for the self-alignment of AI agents with minimal human supervision. Our approach encompasses four stages: first, we use an LLM to generate synthetic prompts, and a topic-guided method to augment the prompt diversity; second, we use a small set of human-written principles for AI models to follow, and guide the LLM through in-context learning from demonstrations (of principles application) to produce helpful, ethical, and reliable responses to user's queries; third, we fine-tune the original LLM with the high-quality self-aligned responses so that the resulting model can generate desirable responses for each query directly without the principle set and the demonstrations anymore; and finally, we offer a refinement step to address the issues of overly-brief or indirect responses. Applying SELF-ALIGN to the LLaMA-65b base language model, we develop an AI assistant named Dromedary. With fewer than 300 lines of human annotations (including < 200 seed prompts, 16 generic principles, and 5 exemplars for in-context learning). Dromedary significantly surpasses the performance of several state-of-the-art AI systems, including Text-Davinci-003 and Alpaca, on benchmark datasets with various settings. Zhiqing Sun, Yikang Shen, Qinhong Zhou, Zhenfang Chen, David D. Cox, Yiming Yang 0002, Chuang Gan 0001 |
NeurIPS | 2 |
| 2023 | Adaptive Online Replanning with Diffusion ModelsabstractDiffusion models have risen a promising approach to data-driven planning, and have demonstrated impressive robotic control, reinforcement learning, and video planning performance. Given an effective planner, an important question to consider is replanning -- when given plans should be regenerated due to both action execution error and external environment changes. Direct plan execution, without replanning, is problematic as errors from individual actions rapidly accumulate and environments are partially observable and stochastic. Simultaneously, replanning at each timestep incurs a substantial computational cost, and may prevent successful task execution, as different generated plans prevent consistent progress to any particular goal. In this paper, we explore how we may effectively replan with diffusion models. We propose a principled approach to determine when to replan, based on the diffusion model's estimated likelihood of existing generated plans. We further present an approach to replan existing trajectories to ensure that new plans follow the same goal state as the original trajectory, which may efficiently bootstrap off previously generated plans. We illustrate how a combination of our proposed additions significantly improves the performance of diffusion planners leading to 38\% gains over past diffusion planning approaches on Maze2D and further enables handling of stochastic and long-horizon robotic control tasks. Yilun Du, Mengdi Xu, Yikang Shen, Wei Xiao 0003, Dit-Yan Yeung, Chuang Gan 0001 |
NeurIPS | 5 |
| 2022 | Phrase-aware Unsupervised Constituency ParsingabstractRecent studies have achieved inspiring success in unsupervised grammar induction using masked language modeling (MLM) as the proxy task.Despite their high accuracy in identifying low-level structures, prior arts tend to struggle in capturing high-level structures like clauses, since the MLM task usually only requires information from local context.In this work, we revisit LM-based constituency parsing from a phrase-centered perspective.Inspired by the natural reading process of human readers, we propose to regularize the parser with phrases extracted by an unsupervised phrase tagger to help the LM model quickly manage low-level structures.For a better understanding of high-level structures, we propose a phrase-guided masking strategy for LM to emphasize more on reconstructing nonphrase words.We show that the initial phrase regularization serves as an effective bootstrap, and phrase-guided masking improves the identification of high-level structures.Experiments on the public benchmark with two different backbone models demonstrate the effectiveness and generality of our method. Xiaotao Gu, Yikang Shen, Jingbo Shang, Jiawei Han 0001 |
ACL (1) | 2 |
| 2022 | Unsupervised Dependency Graph NetworkabstractRecent work has identified properties of pretrained self-attention models that mirror those of dependency parse structures.In particular, some self-attention heads correspond well to individual dependency types.Inspired by these developments, we propose a new competitive mechanism that encourages these attention heads to model different dependency relations.We introduce a new model, the Unsupervised Dependency Graph Network (UDGN), that can induce dependency structures from raw corpora and the masked language modeling task.Experiment results show that UDGN achieves very strong unsupervised dependency parsing performance without gold POS tags and any other external information.The competitive gated heads show a strong correlation with human-annotated dependency types.Furthermore, the UDGN can also achieve competitive performance on masked language modeling and sentence textual similarity tasks 1 . Yikang Shen, Shawn Tan, Alessandro Sordoni, Aaron C. Courville |
ACL (1) | 1 |
| 2022 | Mixture of Attention Heads: Selecting Attention Heads Per TokenabstractMixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing.However, the study of MoE components mostly focused on the feedforward layer in Transformer architecture.This paper proposes the Mixture of Attention Heads (MoA), a new architecture that combines multi-head attention with the MoE mechanism.MoA includes a set of attention heads that each has its own set of parameters.Given an input, a router dynamically selects a subset of k attention heads per token.This conditional computation schema allows MoA to achieve stronger performance than the standard multi-head attention layer.Furthermore, the sparsely gated MoA can easily scale up the number of attention heads and the number of parameters while preserving computational efficiency.In addition to the performance improvements, MoA also automatically differentiates heads' utilities, providing a new perspective to discuss the model's interpretability.We conducted experiments on several important tasks, including Machine Translation and Masked Language Modeling.Experiments have shown promising results on several tasks against strong baselines that involve large and very deep models 1 . Xiaofeng Zhang 0004, Yikang Shen, Wenge Rong, Zhang Xiong 0001 |
EMNLP | 2 |
| 2022 | Prompting Decision Transformer for Few-Shot Policy GeneralizationabstractHuman can leverage prior experience and learn novel tasks from a handful of demonstrations. In contrast to offline meta-reinforcement learning, which aims to achieve quick adaptation through better algorithm design, we investigate the effect of architecture inductive bias on the few-shot learning capability. We propose a Prompt-based Decision Transformer (Prompt-DT), which leverages the sequential modeling ability of the Transformer architecture and the prompt framework to achieve few-shot adaptation in offline RL. We design the trajectory prompt, which contains segments of the few-shot demonstrations, and encodes task-specific information to guide policy generation. Our experiments in five MuJoCo control benchmarks show that Prompt-DT is a strong few-shot learner without any extra finetuning on unseen target tasks. Prompt-DT outperforms its variants and strong meta offline RL baselines by a large margin with a trajectory prompt containing only a few timesteps. Prompt-DT is also robust to prompt length changes and can generalize to out-of-distribution (OOD) environments. Project page: \href{https://mxu34.github.io/PromptDT/}{https://mxu34.github.io/PromptDT/}. Mengdi Xu, Yikang Shen, Ding Zhao, Josh Tenenbaum, Chuang Gan 0001 |
ICML | 2 |
| 2021 | StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingabstractYikang Shen, Yi Tay, Che Zheng, Dara Bahri, Donald Metzler, Aaron Courville. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yikang Shen, Yi Tay, Che Zheng, Dara Bahri, Donald Metzler, Aaron C. Courville |
ACL/IJCNLP (1) | 1 |
| 2021 | Learning Task Decomposition with Ordered Memory Policy Network
Yikang Shen, Aaron C. Courville, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 2 |
| 2021 | Long Range Arena : A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani 0001, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Sebastian Ruder, Donald Metzler |
ICLR | 4 |
| 2021 | Explicitly Modeling Syntax in Language Models with Incremental Parsing and a Dynamic OracleabstractYikang Shen, Shawn Tan, Alessandro Sordoni, Siva Reddy, Aaron Courville. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yikang Shen, Shawn Tan, Alessandro Sordoni, Siva Reddy, Aaron C. Courville |
NAACL-HLT | 1 |
| 2021 | Self-Instantiated Recurrent Units with Dynamic Soft RecursionabstractWhile standard recurrent neural networks explicitly impose a chain structure on different forms of data, they do not have an explicit bias towards recursive self-instantiation where the extent of recursion is dynamic. Given diverse and even growing data modalities (e.g., logic, algorithmic input and output, music, code, images, and language) that can be expressed in sequences and may benefit from more architectural flexibility, we propose the self-instantiated recurrent unit (Self-IRU) with a novel inductive bias towards dynamic soft recursion. On one hand, theSelf-IRU is characterized by recursive self-instantiation via its gating functions, i.e., gating mechanisms of the Self-IRU are controlled by instances of the Self-IRU itself, which are repeatedly invoked in a recursive fashion. On the other hand, the extent of the Self-IRU recursion is controlled by gates whose values are between 0 and 1 and may vary across the temporal dimension of sequences, enabling dynamic soft recursion depth at each time step. The architectural flexibility and effectiveness of our proposed approach are demonstrated across multiple data modalities. For example, the Self-IRU achieves state-of-the-art performance on the logical inference dataset [Bowman et al., 2014] even when comparing with competitive models that have access to ground-truth syntactic information. Aston Zhang, Yi Tay, Yikang Shen, Alvin Chan, Shuai Zhang 0007 |
NeurIPS | 3 |
| 2020 | Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance ApproachabstractIt is commonly believed that knowledge of syntactic structure should improve language modeling.However, effectively and computationally efficiently incorporating syntactic structure into neural language models has been a challenging topic.In this paper, we make use of a multi-task objective, i.e., the models simultaneously predict words as well as ground truth parse trees in a form called "syntactic distances", where information between these two separate objectives shares the same intermediate representation.Experimental results on the Penn Treebank and Chinese Treebank datasets show that when ground truth parse trees are provided as additional training signals, the model is able to achieve lower perplexity and induce trees with better quality. Wenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell, Yoshua Bengio, Yue Zhang 0004 |
ACL | 3 |
| 2019 | Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, Aaron C. Courville |
ICLR | 1 |
| 2019 | Ordered MemoryabstractStack-augmented recurrent neural networks (RNNs) have been of interest to the deep learning community for some time. However, the difficulty of training memory models remains a problem obstructing the widespread use of such models. In this paper, we propose the Ordered Memory architecture. Inspired by Ordered Neurons (Shen et al., 2018), we introduce a new attention-based mechanism and use its cumulative probability to control the writing and erasing operation of the memory. We also introduce a new Gated Recursive Cell to compose lower-level representations into higher-level representation. We demonstrate that our model achieves strong performance on the logical inference task (Bowman et al., 2015) and the ListOps (Nangia and Bowman, 2018) task. We can also interpret the model to retrieve the induced tree structure, and find that these induced structures align with the ground truth. Finally, we evaluate our model on the Stanford Sentiment Treebank tasks (Socher et al., 2013), and find that it performs comparatively with the state-of-the-art methods in the literature. Yikang Shen, Shawn Tan, Seyed Arian Hosseini, Zhouhan Lin, Alessandro Sordoni, Aaron C. Courville |
NeurIPS | 1 |
| 2018 | Straight to the Tree: Constituency Parsing with Neural Syntactic DistanceabstractYikang Shen, Zhouhan Lin, Athul Paul Jacob, Alessandro Sordoni, Aaron Courville, Yoshua Bengio. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Yikang Shen, Zhouhan Lin, Athul Paul Jacob, Alessandro Sordoni, Aaron C. Courville, Yoshua Bengio |
ACL (1) | 1 |
| 2018 | BanditSum: Extractive Summarization as a Contextual BanditabstractIn this work, we propose a novel method for training neural networks to perform singledocument extractive summarization without heuristically-generated extractive labels.We call our approach BANDITSUM as it treats extractive summarization as a contextual bandit (CB) problem, where the model receives a document to summarize (the context), and chooses a sequence of sentences to include in the summary (the action).A policy gradient reinforcement learning algorithm is used to train the model to select sequences of sentences that maximize ROUGE score.We perform a series of experiments demonstrating that BANDITSUM is able to achieve ROUGE scores that are better than or comparable to the state-of-the-art for extractive summarization, and converges using significantly fewer update steps than competing approaches.In addition, we show empirically that BANDIT-SUM performs significantly better than competing approaches when good summary sentences appear late in the source document. Yue Dong 0002, Yikang Shen, Eric Crawford, Herke van Hoof, Jackie Chi Kit Cheung |
EMNLP | 2 |
| 2018 | Neural Language Modeling by Jointly Learning Syntax and Lexicon
Yikang Shen, Zhouhan Lin, Chin-Wei Huang, Aaron C. Courville |
ICLR (Poster) | 1 |
| 2018 | Biological Event Trigger Identification with Noise Contrastive EstimationabstractBiological Event Extraction is an important task towards the goal of extracting biomedical knowledge from the scientific publications by capturing biomedical entities and their complex relations from the texts. As a crucial step in event extraction, event trigger identification, assigning words with suitable trigger category, has recently attracted substantial attention. As triggers are scattered in large corpus, traditional linguistic parsers are hard to generate syntactic features from them. Thereby, trigger sparsity problem restricts the model's learning process and becomes one of the main hinder in trigger identification. In this paper, we employ Noise Contrastive Estimation with Multi-Layer Perceptron model for solving triggers' sparsity problem. Meanwhile, in the light of recent advance in word distributed representation, word-embedding feature generated by language model is utilized for semantic and syntactic information extraction. Finally, experimental study on commonly used MLEE dataset against baseline methods has demonstrated its promising result. Nan Jiang 0010, Wenge Rong, Yifan Nie, Yikang Shen, Zhang Xiong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2017 | Word Embedding Based Correlation Model for Question/Answer MatchingabstractThe large scale of Q&A archives accumulated in community based question answering (CQA) servivces are important information and knowledge resource on the web. Question and answer matching task has been attached much importance to for its ability to reuse knowledge stored in these systems: it can be useful in enhancing user experience with recurrent questions. In this paper, a Word Embedding based Correlation (WEC) model is proposed by integrating advantages of both the translation model and word embedding. Given a random pair of words, WEC can score their co-occurrence probability in Q&A pairs, while it can also leverage the continuity and smoothness of continuous space word representation to deal with new pairs of words that are rare in the training parallel text. An experimental study on Yahoo! Answers dataset and Baidu Zhidao dataset shows this new method's promising potential. Yikang Shen, Wenge Rong, Nan Jiang 0010, Baolin Peng, Jie Tang 0001, Zhang Xiong 0001 |
AAAI | 1 |
| 2017 | Exploration of Tree-based Hierarchical Softmax for Recurrent Language ModelsabstractRecently, variants of neural networks for computational linguistics have been proposed and successfully applied to neural language modeling and neural machine translation. These neural models can leverage knowledge from massive corpora but they are extremely slow as they predict candidate words from a large vocabulary during training and inference. As an alternative to gradient approximation and softmax with class decomposition, we explore the tree-based hierarchical softmax method and reform its architecture, making it compatible with modern GPUs and introducing a compact tree-based loss function. When combined with several word hierarchical clustering algorithms, improved performance is achieved in language modelling task with intrinsic evaluation criterions on PTB, WikiText-2 and WikiText-103 datasets. Nan Jiang 0010, Wenge Rong, Min Gao 0001, Yikang Shen, Zhang Xiong 0001 |
IJCAI | 4 |
| 2016 | Convolutional Neural Network based sentiment analysis using Adaboost combinationabstractSentimental polarity detection has long been a hot task in natural language processing since its applications range from product feedback analysis to user statement understanding. Recently a lot of machine learning approaches have been proposed in the literature, e.g., SVM, Naive Bayes, recursive neural network, auto-encoders and etc. Among these different models, Convolutional Neural Network (CNN) architecture have also demonstrated profound efficiency in NLP tasks including sentiment classification. In CNN, the width of convolutional filter functions alike number N in N-grams model. Thus, different filter lengths may influence the performance of CNN classifier. In this paper, we want to study the possibility of leveraging the contribution of different filter lengths and grasp their potential in the final polarity of the sentence. We then use Adaboost to combine different classifiers with respective filter sizes. The experimental study on commonly used datasets has shown its potential in identifying the different roles of specific N-grams in a sentence respectively and merging their contribution in a weighted classifier. Yazhi Gao, Wenge Rong, Yikang Shen, Zhang Xiong 0001 |
IJCNN | 3 |
| 2016 | Multidimensional scaling based knowledge provision for new questions in community Question Answering systemsabstractCommunity-based Question Answering (CQA) sites have become popular since they allow users to get answers to complex, detailed and personal question from other users directly. However, since answering a question depends on the ability and willingness of other users to address the askers' real needs, a significant fraction of the questions remain unanswered. To decrease the unanswered question rate and then improve the user experience, in this paper, a multidimensional scaling (MDS) based data reorganization method is proposed. By using this method, the CQA system can predict the askers' intention and accordingly provide related previous question/answer pairs to help them find useful information. The method has been evaluated on an off-line dataset extracted from Baidu Zhidao and the result has shown its promising potential in knowledge management in CQA systems. Siqi Xiang, Wenge Rong, Yikang Shen, Yuanxin Ouyang, Zhang Xiong 0001 |
IJCNN | 3 |
| 2015 | Question/Answer Matching for CQA System via Combining Lexical and Sequential InformationabstractCommunity-based Question Answering (CQA) has become popular in knowledge sharing sites since it allows users to get answers to complex, detailed, and personal questions directly from other users. Large archives of historical questions and associated answers have been accumulated. Retrieving relevant historical answers that best match a question is an essential component of a CQA service. Most state of the art approaches are based on bag-of-words models, which have been proven successful in a range of text matching tasks, but are insufficient for capturing the important word sequence information in short text matching. In this paper, a new architecture is proposed to more effectively model the complicated matching relations between questions and answers. It utilises a similarity matrix which contains both lexical and sequential information. Afterwards the information is put into a deep architecture to find potentially suitable answers. The experimental study shows its potential in improving matching accuracy of question and answer. Yikang Shen, Wenge Rong, Yuanxin Ouyang, Zhang Xiong 0001 |
AAAI | 1 |
| 2014 | Choosing the Best Auto-Encoder-Based Bagging Classifier: An Empirical Study
Yifan Nie, Wenge Rong, Yikang Shen, Chao Li 0001, Zhang Xiong 0001 |
ICONIP (1) | 3 |