EDBT 2026 Demo / reviewers in the wild / expert
Chen-Yu Lee
dblp:04/656
· DBLP profile ↗
49ranked-venue papers
11as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When One LLM Drools, Multi-LLM Collaboration RulesabstractShangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang, Weijia Shi, Yike Wang, Shannon Zejiang Shen, Xiaochuang Han, Hunter Lang, Chen-Yu Lee, Tomas Pfister, Yejin Choi, Yulia Tsvetkov. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shangbin Feng, Wenxuan Ding 0001, Alisa Liu, Zifeng Wang 0002, Yike Wang 0002, Shannon Shen 0001, Xiaochuang Han, Hunter Lang, Chen-Yu Lee, Tomas Pfister, Yejin Choi 0001, Yulia Tsvetkov |
ACL (1) | 10 |
| 2025 | In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue AgentsabstractLarge Language Models (LLMs) have made significant progress in open-ended dialogue, yet their inability to retain and retrieve relevant information from long-term interactions limits their effectiveness in applications requiring sustained personalization. External memory mechanisms have been proposed to address this limitation, enabling LLMs to maintain conversational continuity. However, existing approaches struggle with two key challenges. First, rigid memory granularity fails to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations. Second, fixed retrieval mechanisms cannot adapt to diverse dialogue contexts and user interaction patterns. In this work, we propose Reflective Memory Management (RMM), a novel mechanism for long-term dialogue agents, integrating forward- and backward-looking reflections: (1) Prospective Reflection, which dynamically summarizes interactions across granularities—utterances, turns, and sessions—into a personalized memory bank for effective future retrieval, and (2) Retrospective Reflection, which iteratively refines the retrieval in an online reinforcement learning (RL) manner based on LLMs’ cited evidence. Experiments show that RMM demonstrates consistent improvement across various metrics and benchmarks. For example, RMM shows more than 10% accuracy improvement over the baseline without memory management on the LongMemEval dataset. Zhen Tan 0001, Jun Yan 0001, I-Hung Hsu, Rujun Han, Zifeng Wang 0002, Long T. Le, Yiwen Song, Yanfei Chen, Hamid Palangi, Anand Rajan Iyer, Tianlong Chen 0001, Huan Liu 0001, Chen-Yu Lee, Tomas Pfister |
ACL (1) | 14 |
| 2025 | Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph TranslationabstractFan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan, Ke Jiang, Yanfei Chen, Jindong Gu, Long Le, Kai-Wei Chang, Chen-Yu Lee, Hamid Palangi, Tomas Pfister. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fan Yin, Zifeng Wang 0002, I-Hung Hsu, Jun Yan 0001, Yanfei Chen, Jindong Gu, Long T. Le, Kai-Wei Chang 0001, Chen-Yu Lee, Hamid Palangi, Tomas Pfister |
ACL (1) | 10 |
| 2025 | PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem SolvingabstractMihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mihir Parmar, Palash Goyal, Yanfei Chen, Long T. Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang 0002, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi |
EMNLP | 12 |
| 2025 | Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingabstractRetrieval augmented generation (RAG) combines the generative abilities of large language models (LLMs) with external knowledge sources to provide more accurate and up-to-date responses. Recent RAG advancements focus on improving retrieval outcomes through iterative LLM refinement or self-critique capabilities acquired through additional instruction tuning of LLMs. In this work, we introduce Speculative RAG - a framework that leverages a larger generalist LM to efficiently verify multiple RAG drafts produced in parallel by a smaller, distilled specialist LM. Each draft is generated from a distinct subset of retrieved documents, offering diverse perspectives on the evidence while reducing input token counts per draft. This approach enhances comprehension of each subset and mitigates potential position bias over long context. Our method accelerates RAG by delegating drafting to the smaller specialist LM, with the larger generalist LM performing a single verification pass over the drafts. Extensive experiments demonstrate that Speculative RAG achieves state-of-the-art performance with reduced latency on TriviaQA, MuSiQue, PopQA, PubHealth, and ARC-Challenge benchmarks. It notably enhances accuracy by up to 12.97% while reducing latency by 50.83% compared to conventional RAG systems on PubHealth. Zilong Wang 0002, Zifeng Wang 0002, Long T. Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang 0001, Anush Mattapalli, Ankur Taly, Jingbo Shang, Chen-Yu Lee, Tomas Pfister |
ICLR | 11 |
| 2025 | Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved SamplingabstractRecent advances in knowledge distillation (KD) have enabled smaller student models to approach the performance of larger teacher models. However, popular methods such as supervised KD and on-policy KD, are adversely impacted by the knowledge gaps between teacher-student in practical scenarios. Supervised KD suffers from a distribution mismatch between training with a static dataset and inference over final student-generated outputs. Conversely, on-policy KD, which uses student-generated samples for training, can suffer from low-quality training examples with which teacher models are not familiar, resulting in inaccurate teacher feedback. To address these limitations, we introduce Speculative Knowledge Distillation (SKD), a novel approach that leverages cooperation between student and teacher models to generate high-quality training data on-the-fly while aligning with the student's inference-time distribution. In SKD, the student proposes tokens, and the teacher replaces poorly ranked ones based on its own distribution, transferring high-quality knowledge adaptively. We evaluate SKD on various text generation tasks, including translation, summarization, math, and instruction following, and show that SKD consistently outperforms existing KD methods across different domains, data sizes, and model initialization strategies. Wenda Xu, Rujun Han, Zifeng Wang 0002, Long T. Le, Dhruv Madeka, Lei Li 0005, William Yang Wang, Rishabh Agarwal, Chen-Yu Lee, Tomas Pfister |
ICLR | 9 |
| 2025 | Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm IntelligenceabstractWe propose Model Swarms, a collaborative search algorithm to adapt LLMs via swarm intelligence, the collective behavior guiding individual systems. Specifically, Model Swarms starts with a pool of LLM experts and a utility function. Guided by the best-found checkpoints across models, diverse LLM experts collaboratively move in the weight space and optimize a utility function representing model adaptation objectives. Compared to existing model composition approaches, Model Swarms offers tuning-free model adaptation, works in low-data regimes with as few as 200 examples, and does not require assumptions about specific experts in the swarm or how they should be composed. Extensive experiments demonstrate that Model Swarms could flexibly adapt LLM experts to a single task, multi-task domains, reward models, as well as diverse human interests, improving over 12 model composition baselines by up to 21.0% across tasks and contexts. Further analysis reveals that LLM experts discover previously unseen capabilities in initial checkpoints and that Model Swarms enable the weak-to-strong transition of experts through the collaborative search process. Shangbin Feng, Zifeng Wang 0002, Yike Wang 0002, Sayna Ebrahimi, Hamid Palangi, Lesly Miculicich, Achin Kulshrestha, Nathalie Rauschmayr, Yejin Choi 0001, Yulia Tsvetkov, Chen-Yu Lee, Tomas Pfister |
ICML | 11 |
| 2025 | AI for Supply Chain: Today and FutureabstractA supply chain is the network of entities and processes involved in the production and distribution of a commodity. Supply chains are a critical backbone across industries like retail, manufacturing, healthcare, and automotive, driving everything from product availability to operational efficiency and customer satisfaction. Modern supply chains are 1) non-cooperative, functioning as fragmented systems where isolated technologies solve individual problems without integration, and 2) unadaptable, failing to adjust to real-time data and uncertainties, like demand fluctuations and regulatory changes. As a result, frequent manual overrides are required, as even small errors can lead to significant financial losses, strained customer relationships, or reputational damage. Modern Artificial Intelligence (AI) advancements offer great potential to unify fragmented supply chains into a seamless, adaptive system while enhancing automation, decision-making, and transparency. In this workshop, we will examine critical supply chain challenges and demonstrate how AI can provide accurate, efficient, and scalable solutions. Ranak Roy Chowdhury, Yan Liu 0002, Huiming Qu, Qingsong Wen, Chen-Yu Lee, Narendra Agrawal, Alexis Roos |
KDD (2) | 5 |
| 2025 | Reverse Thinking Makes LLMs Stronger ReasonersabstractJustin Chen, Zifeng Wang, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Justin Chih-Yao Chen, Zifeng Wang 0002, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long T. Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister |
NAACL (Long Papers) | 10 |
| 2025 | Where is the answer? An empirical study of positional bias for parametric knowledge extraction in language modelabstractKuniaki Saito, Chen-Yu Lee, Kihyuk Sohn, Yoshitaka Ushiku. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kuniaki Saito, Chen-Yu Lee, Kihyuk Sohn, Yoshitaka Ushiku |
NAACL (Long Papers) | 2 |
| 2025 | Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsabstractWe propose Heterogeneous Swarms, an algorithm to design multi-LLM systems by jointly optimizing model roles and weights. We represent multi-LLM systems as directed acyclic graphs (DAGs) of LLMs with topological message passing for collaborative generation. Given a pool of LLM experts and a utility function, Heterogeneous Swarms employs two iterative steps: role-step and weight-step. For role-step, we interpret model roles as learning a DAG that specifies the flow of inputs and outputs between LLMs. Starting from a swarm of random continuous adjacency matrices, we decode them into discrete DAGs, call the LLMs in topological order, evaluate on the utility function (e.g. accuracy on a task), and optimize the adjacency matrices with particle swarm optimization based on the utility score. For weight-step, we assess the contribution of individual LLMs in the multi-LLM systems and optimize model weights with swarm intelligence. We propose JFK-score to quantify the individual contribution of each LLM in the best-found DAG of the role-step, then optimize model weights with particle swarm optimization based on the JFK-score. Experiments demonstrate that Heterogeneous Swarms outperforms 17 role- and/or weight-based baselines by 18.5% on average across 12 tasks. Further analysis reveals that Heterogeneous Swarms discovers multi-LLM systems with heterogeneous model roles and substantial collaborative gains, and benefits from the diversity of language models. Shangbin Feng, Zifeng Wang 0002, Palash Goyal, Yike Wang 0002, Huang Xia, Hamid Palangi, Luke Zettlemoyer, Yulia Tsvetkov, Chen-Yu Lee, Tomas Pfister |
NeurIPS | 10 |
| 2024 | Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingabstractTable-based reasoning with large language models (LLMs) is a promising direction to tackle many table understanding tasks, such as table-based question answering and fact verification. Compared with generic reasoning, table-based reasoning requires the extraction of underlying semantics from both free-form questions and semi-structured tabular data. Chain-of-Thought and its similar approaches incorporate the reasoning chain in the form of textual context, but it is still an open question how to effectively leverage tabular data in the reasoning chain. We propose the Chain-of-Table framework, where tabular data is explicitly used in the reasoning chain as a proxy for intermediate thoughts. Specifically, we guide LLMs using in-context learning to iteratively generate operations and update the table to represent a tabular reasoning chain. LLMs can therefore dynamically plan the next operation based on the results of the previous ones. This continuous evolution of the table forms a chain, showing the reasoning process for a given tabular problem. The chain carries structured information of the intermediate results, enabling more accurate and reliable predictions. Chain-of-Table achieves new state-of-the-art performance on WikiTQ, FeTaQA, and TabFact benchmarks across multiple LLM choices. Zilong Wang 0002, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang 0002, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, Tomas Pfister |
ICLR | 10 |
| 2024 | TableRAG: Million-Token Table Understanding with Language ModelsabstractRecent advancements in language models (LMs) have notably enhanced their ability to reason with tabular data, primarily through program-aided mechanisms that manipulate and analyze tables.
However, these methods often require the entire table as input, leading to scalability challenges due to the positional bias or context length constraints.
In response to these challenges, we introduce TableRAG, a Retrieval-Augmented Generation (RAG) framework specifically designed for LM-based table understanding.
TableRAG leverages query expansion combined with schema and cell retrieval to pinpoint crucial information before providing it to the LMs.
This enables more efficient data encoding and precise retrieval, significantly reducing prompt lengths and mitigating information loss.
We have developed two new million-token benchmarks from the Arcade and BIRD-SQL datasets to thoroughly evaluate TableRAG's effectiveness at scale.
Our results demonstrate that TableRAG's retrieval design achieves the highest retrieval quality, leading to the new state-of-the-art performance on large-scale table understanding. Si-An Chen, Lesly Miculicich, Julian Martin Eisenschlos, Zifeng Wang 0002, Zilong Wang 0002, Yanfei Chen, Yasuhisa Fujii, Hsuan-Tien Lin, Chen-Yu Lee, Tomas Pfister |
NeurIPS | 9 |
| 2023 | Neural Spline Search for Quantile Probabilistic ModelingabstractAccurate estimation of output quantiles is crucial in many use cases, where it is desired to model the range of possibility. Modeling target distribution at arbitrary quantile levels and at arbitrary input attribute levels are important to offer a comprehensive picture of the data, and requires the quantile function to be expressive enough. The quantile function describing the target distribution using quantile levels is critical for quantile regression. Although various parametric forms for the distributions (that the quantile function specifies) can be adopted, an everlasting problem is selecting the most appropriate one that can properly approximate the data distributions. In this paper, we propose a non-parametric and data-driven approach, Neural Spline Search (NSS), to represent the observed data distribution without parametric assumptions. NSS is flexible and expressive for modeling data distributions by transforming the inputs with a series of monotonic spline regressions guided by symbolic operators. We demonstrate that NSS outperforms previous methods on synthetic, real-world regression and time-series forecasting tasks. Ruoxi Sun 0002, Chun-Liang Li, Sercan Ö. Arik, Michael Dusenberry, Chen-Yu Lee, Tomas Pfister |
AAAI | 5 |
| 2023 | FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information ExtractionabstractChen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister |
ACL (1) | 1 |
| 2023 | Multimodal Prompting with Missing Modalities for Visual RecognitionabstractIn this paper, we tackle two challenges in multimodal learning for visual recognition: 1) when missing-modality occurs either during training or testing in real-world situations; and 2) when the computation resources are not available to finetune on heavy transformer models. To this end, we propose to utilize prompt learning and mitigate the above two challenges together. Specifically, our modality-missing-aware prompts can be plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 1% learnable parameters compared to training the entire model. We further explore the effect of different prompt configurations and analyze the robustness to missing modality. Extensive experiments are conducted to show the effectiveness of our prompt learning framework that improves the performance under various missing-modality cases, while alleviating the requirement of heavy model retraining. Code is available.11https://github.com/YiLunLee/missing_aware_prompts Yi-Lun Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu, Chen-Yu Lee |
CVPR | 4 |
| 2023 | Prefix Conditioning Unifies Language and Label SupervisionabstractPretraining visual models on web-scale image-caption datasets has recently emerged as a powerful alternative to traditional pretraining on image classification data. Image-caption datasets are more “opendomain ”, containing broader scene types and vocabulary words, and result in models that have strong performance in fewand zero-shot recognition tasks. However large-scale classification datasets can provide fine-grained categories with a balanced label distribution. In this work, we study a pretraining strategy that uses both classification and caption datasets to unite their complementary benefits. First, we show that naively unifying the datasets results in sub-optimal performance in downstream zero-shot recognition tasks, as the model is affected by dataset bias: the coverage of image domains and vocabulary words is different in each dataset. We address this problem with novel Prefix Conditioning, a simple yet effective method that helps disentangle dataset biases from visual concepts. This is done by intro-ducing prefix tokens that inform the language encoder of the input data type (e.g., classification vs caption) at training time. Our approach allows the language encoder to learn from both datasets while also tailoring feature extraction to each dataset. Prefix conditioning is generic and can be easily integrated into existing VL pretraining objectives, such as CLIP or UniCL. In experiments, we show that it improves zero-shot image recognition and robustness to image-level distribution shift. Kuniaki Saito, Kihyuk Sohn, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, Tomas Pfister |
CVPR | 5 |
| 2023 | Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalabstractIn Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using labeled triplets consisting of the query image, text specification, and the target image. Labeling such triplets is expensive and hinders broad applicability of CIR. In this work, we propose to study an important task, Zero-Shot Composed Image Retrieval (ZS-CIR), whose goal is to build a CIR model without requiring labeled triplets for training. To this end, we propose a novel method, called Pic2Word, that requires only weakly labeled image-caption pairs and unlabeled image datasets to train. Unlike existing supervised CIR models, our model trained on weakly labeled or unlabeled datasets shows strong generalization across diverse ZS-CIR tasks, e.g., attribute editing, object composition, and domain conversion. Our approach outperforms several supervised CIR methods on the common CIR benchmark, CIRR and Fashion-IQ. Code will be made publicly available at https://github.com/google-research/composed_image_retrieval Kuniaki Saito, Kihyuk Sohn, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, Tomas Pfister |
CVPR | 5 |
| 2023 | VRDU: A Benchmark for Visually-rich Document UnderstandingabstractUnderstanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achieved impressive results, we find that existing benchmarks do not reflect the complexity of real documents seen in industry. In this work, we identify the desiderata for a more comprehensive benchmark and propose one we call Visually Rich Document Understanding (VRDU). VRDU contains two datasets that represent several challenges: rich schema including diverse data types as well as hierarchical entities, complex templates including tables and multi-column layouts, and diversity of different layouts (templates) within a single document type. We design few-shot and conventional experiment settings along with a carefully designed matching algorithm to evaluate extraction results. We report the performance of strong baselines and offer three observations: (1) generalizing to new document templates is still very challenging, (2) few-shot performance has a lot of headroom, and (3) models struggle with hierarchical fields such as line-items in an invoice. We plan to open source the benchmark and the evaluation toolkit. We hope this helps the community make progress on these challenging tasks in extracting structured data from visually rich documents. Zilong Wang 0002, Yichao Zhou 0001, Wei Wei 0019, Chen-Yu Lee, Sandeep Tata |
KDD | 4 |
| 2023 | Data Efficient Incremental Learning via Attentive Knowledge ReplayabstractClass-incremental learning (CIL) tackles the problem of continuously optimizing a classification model to support growing number of classes, where the data of novel classes arrive in streams. Recent works propose to use representative exemplars of learnt classes, and replay the knowledge of them afterward under certain memory constraints. However, training on a fixed set of exemplars with an imbalanced proportion to the new data leads to strong biases in the trained models. In this paper, we propose an attentive knowledge replay framework to refresh the knowledge of previously learnt classes during incremental learning, which generates virtual training samples by blending between pairs of data. Particularly, we design an attention module that learns to predict the adaptive blending weights in accordance with their relative importance to the overall objective, where the importance is derived from the change of the image features over incremental phases. Our strategy of attentive knowledge replay encourages the model to learn smoother decision boundaries and thus improves its generalization beyond memorizing the exemplars. We validate our design in a standard class-incremental learning setup and demonstrate its flexibility in various settings. Yi-Lun Lee, Dian-Shan Chen, Chen-Yu Lee, Yi-Hsuan Tsai, Walon Wei-Chen Chiu |
SMC | 3 |
| 2023 | Unifying Distribution Alignment as a Loss for Imbalanced Semi-supervised LearningabstractWhile remarkable progress has been made in imbalanced supervised learning, less attention has been given to the setting of imbalanced semi-supervised learning (SSL) where not only are few labeled data provided, but the underlying data distribution can be severely imbalanced. Recent work requires both complicated sampling strategies of pseudo-labeled unlabeled data and distribution alignment of the pseudo-label distribution to accommodate this imbalance. We present a novel approach that relies only on a form of a distribution alignment but no sampling strategy where rather than aligning the pseudo-labels during inference, we move the distribution alignment component into the respective cross entropy loss computations for both the supervised and unsupervised losses. This alignment compensates for both imbalance in the data and the eventual distributional shift present during evaluation. Altogether, this provides a unified strategy that offers both significantly reduced training requirements and improved performance across both low and richly labeled regimes and over varying degrees of imbalance. In experiments, we validate the efficacy of our method on SSL variants of CIFAR10-LT, CIFAR100-LT, and ImageNet-127. On ImageNet-127, our method shows 1.6% accuracy improvement over CReST with an 80% training time reduction and is competitive with other SOTA methods. Code is available at https://github.com/google-research/crest Justin Lazarow, Kihyuk Sohn, Chen-Yu Lee, Chun-Liang Li, Tomas Pfister |
WACV | 3 |
| 2023 | Anomaly Clustering: Grouping Images into Coherent Clusters of Anomaly TypesabstractWe study anomaly clustering, grouping data into coherent clusters of anomaly types. This is different from anomaly detection that aims to divide anomalies from normal data. Unlike object-centered image clustering, anomaly clustering is particularly challenging as anomalous patterns are subtle and local. We present a simple yet effective clustering framework using a patch-based pretrained deep embeddings and off-the-shelf clustering methods. We define a distance function between images, each of which is represented as a bag of embeddings, by the Euclidean distance between weighted averaged embeddings. The weight defines the importance of instances (i.e., patch embeddings) in the bag, which may highlight defective regions. We compute weights in an unsupervised way or in a semi-supervised way when labeled normal data is available. Extensive experimental studies show the effectiveness of the proposed clustering framework along with a novel distance function upon existing multiple instance or deep clustering frameworks. Overall, our framework achieves 0.451 and 0.674 normalized mutual information scores on MVTec object and texture categories and further improve with a few labeled normal data (0.577, 0.669), far exceeding the baselines (0.244, 0.273) or state-of-the-art deep clustering methods (0.176, 0.277). Kihyuk Sohn, Jinsung Yoon, Chun-Liang Li, Chen-Yu Lee, Tomas Pfister |
WACV | 4 |
| 2022 | Learning from Weakly-Labeled Web Videos via Exploring Sub-conceptsabstractLearning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet. However, for video action recognition, the action of interest might only exist in arbitrary clips of untrimmed web videos, resulting in high label noises in the temporal space. To address this challenge, we introduce a new method for pre-training video action recognition models using queried web videos. Instead of trying to filter out potential noises, we propose to provide fine-grained supervision signals by defining the concept of Sub-Pseudo Label (SPL). Specifically, SPL spans out a new set of meaningful "middle ground" label space constructed by extrapolating the original weak labels during video querying and the prior knowledge distilled from a teacher model. Consequently, SPL provides enriched supervision for video models to learn better representations and improves data utilization efficiency of untrimmed videos. We validate the effectiveness of our method on four video action recognition datasets and a weakly-labeled image dataset. Experiments show that SPL outperforms several existing pre-training strategies and the learned representations lead to competitive results on several benchmarks. Guanhang Wu, Xuehan Xiong, Chen-Yu Lee, Zhichao Lu, Yun Fu 0001, Tomas Pfister |
AAAI | 5 |
| 2022 | FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionabstractChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, Tomas Pfister. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, Tomas Pfister |
ACL (1) | 1 |
| 2022 | Learning to Prompt for Continual LearningabstractThe mainstream paradigm behind continual learning has been to adapt the model parameters to non-stationary data distributions, where catastrophic forgetting is the central challenge. Typical methods rely on a rehearsal buffer or known task identity at test time to retrieve learned knowl-edge and address forgetting, while this work presents a new paradigm for continual learning that aims to train a more succinct memory system without accessing task identity at test time. Our method learns to dynamically prompt (L2P) a pre-trained model to learn tasks sequen-tially under different task transitions. In our proposed framework, prompts are small learnable parameters, which are maintained in a memory space. The objective is to optimize prompts to instruct the model prediction and ex-plicitly manage task-invariant and task-specific knowledge while maintaining model plasticity. We conduct comprehen-sive experiments under popular image classification bench-marks with different challenging continual learning set-tings, where L2P consistently outperforms prior state-of-the-art methods. Surprisingly, L2P achieves competitive results against rehearsal-based methods even without a re-hearsal buffer and is directly applicable to challenging task-agnostic continual learning. Source code is available at https://github.com/google-research/12p. Zifeng Wang 0002, Chen-Yu Lee, Han Zhang 0010, Ruoxi Sun 0002, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer G. Dy, Tomas Pfister |
CVPR | 3 |
| 2022 | DualPrompt: Complementary Prompting for Rehearsal-Free Continual Learning
Zifeng Wang 0002, Sayna Ebrahimi, Ruoxi Sun 0002, Han Zhang 0010, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer G. Dy, Tomas Pfister |
ECCV (26) | 6 |
| 2020 | Learning to Branch for Multi-Task LearningabstractTraining multiple tasks jointly in one deep network yields reduced latency during inference and better performance over the single-task counterpart by sharing certain layers of a network. However, over-sharing a network could erroneously enforce over-generalization, causing negative knowledge transfer across tasks. Prior works rely on human intuition or pre-computed task relatedness scores for ad hoc branching structures. They provide sub-optimal end results and often require huge efforts for the trial-and-error process. In this work, we present an automated multi-task learning algorithm that learns where to share or branch within a network, designing an effective network topology that is directly optimized for multiple objectives across tasks. Specifically, we propose a novel tree-structured design space that casts a tree branching operation as a gumbel-softmax sampling procedure. This enables differentiable network splitting that is end-to-end trainable. We validate the proposed method on controlled synthetic data, CelebA, and Taskonomy. Pengsheng Guo, Chen-Yu Lee, Daniel Ulbricht |
ICML | 2 |
| 2019 | Sliced Wasserstein Discrepancy for Unsupervised Domain AdaptationabstractIn this work, we connect two distinct concepts for unsupervised domain adaptation: feature distribution alignment between domains by utilizing the task-specific decision boundary and the Wasserstein metric. Our proposed sliced Wasserstein discrepancy (SWD) is designed to capture the natural notion of dissimilarity between the outputs of task-specific classifiers. It provides a geometrically meaningful guidance to detect target samples that are far from the support of the source and enables efficient distribution alignment in an end-to-end trainable fashion. In the experiments, we validate the effectiveness and genericness of our method on digit and sign recognition, image classification, semantic segmentation, and object detection. Chen-Yu Lee, Tanmay Batra, Mohammad Haris Baig, Daniel Ulbricht |
CVPR | 1 |
| 2018 | GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask NetworksabstractDeep multitask networks, in which one neural network produces multiple predictive outputs, can offer better speed and performance than their single-task counterparts but are challenging to train properly. We present a gradient normalization (GradNorm) algorithm that automatically balances training in deep multitask models by dynamically tuning gradient magnitudes. We show that for various network architectures, for both regression and classification tasks, and on both synthetic and real datasets, GradNorm improves accuracy and reduces overfitting across multiple tasks when compared to single-task networks, static baselines, and other adaptive multitask loss balancing techniques. GradNorm also matches or surpasses the performance of exhaustive grid search methods, despite only involving a single asymmetry hyperparameter $\alpha$. Thus, what was once a tedious search process that incurred exponentially more compute for each task added can now be accomplished within a few training runs, irrespective of the number of tasks. Ultimately, we will demonstrate that gradient manipulation affords us great control over the training dynamics of multitask networks and may be one of the keys to unlocking the potential of multitask learning. Zhao Chen 0006, Vijay Badrinarayanan, Chen-Yu Lee, Andrew Rabinovich |
ICML | 3 |
| 2018 | Generalizing Pooling Functions in CNNs: Mixed, Gated, and TreeabstractIn this paper, we seek to improve deep neural networks by generalizing the pooling operations that play a central role in the current architectures. We pursue a careful exploration of approaches to allow pooling to learn and to adapt to complex and variable patterns. The two primary directions lie in: (1) learning a pooling function via (two strategies of) combining of max and average pooling, and (2) learning a pooling function in the form of a tree-structured fusion of pooling filters that are themselves learned. In our experiments every generalized pooling operation we explore improves performance when used in place of average or max pooling. We experimentally demonstrate that the proposed pooling operations provide a boost in invariance properties relative to conventional pooling and set the state of the art on several widely adopted benchmark datasets. These benefits come with only a light increase in computational overhead during training (ranging from additional 5 to 15 percent in time complexity) and a very modest increase in the number of model parameters (e.g., additional 1, 9, and 27 parameters for mixed, gated, and 2-level tree pooling operators, respectively). To gain more insights about our proposed pooling methods, we also visualize the learned pooling masks and the embeddings of the internal feature responses for different pooling operations. Our proposed pooling operations are easy to implement and can be applied within various deep neural network architectures. Chen-Yu Lee, Patrick W. Gallagher, Zhuowen Tu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | RoomNet: End-to-End Room Layout EstimationabstractThis paper focuses on the task of room layout estimation from a monocular RGB image. Prior works break the problem into two sub-tasks: semantic segmentation of floor, walls, ceiling to produce layout hypotheses, followed by an iterative optimization step to rank these hypotheses. In contrast, we adopt a more direct formulation of this problem as one of estimating an ordered set of room layout keypoints. The room layout and the corresponding segmentation is completely specified given the locations of these ordered keypoints. We predict the locations of the room layout keypoints using RoomNet, an end-to-end trainable encoder-decoder network. On the challenging benchmark datasets Hedau and LSUN, we achieve state-of-the-art performance along with 200× to 600× speedup compared to the most recent work. Additionally, we present optional extensions to the RoomNet architecture such as including recurrent computations and memory units to refine the keypoint locations under the same parametric capacity. Chen-Yu Lee, Vijay Badrinarayanan, Tomasz Malisiewicz, Andrew Rabinovich |
ICCV | 1 |
| 2016 | Generalizing Pooling Functions in Convolutional Neural Networks: Mixed, Gated, and TreeabstractWe seek to improve deep neural networks by generalizing the pooling operations that play a central role in current architectures. We pursue a careful exploration of approaches to allow pooling to learn and to adapt to complex and variable patterns. The two primary directions lie in (1) learning a pooling function via (two strategies of) combining of max and average pooling, and (2) learning a pooling function in the form of a tree-structured fusion of pooling filters that are themselves learned. In our experiments every generalized pooling operation we explore improves performance when used in place of average or max pooling. We experimentally demonstrate that the proposed pooling operations provide a boost in invariance properties relative to conventional pooling and set the state of the art on several widely adopted benchmark datasets; they are also easy to implement, and can be applied within various deep neural network architectures. These benefits come with only a light increase in computational overhead during training and a very modest increase in the number of model parameters. Chen-Yu Lee, Patrick W. Gallagher, Zhuowen Tu |
AISTATS | 1 |
| 2016 | Recursive Recurrent Nets with Attention Modeling for OCR in the WildabstractWe present recursive recurrent neural networks with attention modeling (R2AM) for lexicon-free optical character recognition in natural scene images. The primary advantages of the proposed method are: (1) use of recursive convolutional neural networks (CNNs), which allow for parametrically efficient and effective image feature extraction, (2) an implicitly learned character-level language model, embodied in a recurrent neural network which avoids the need to use N-grams, and (3) the use of a soft-attention mechanism, allowing the model to selectively exploit image features in a coordinated way, and allowing for end-to-end training within a standard backpropagation framework. We validate our method with state-of-the-art performance on challenging benchmark datasets: Street View Text, IIIT5k, ICDAR and Synth90k. Chen-Yu Lee, Simon Osindero |
CVPR | 1 |
| 2015 | Deeply-Supervised NetsabstractWe propose deeply-supervised nets (DSN), a method that simultaneously minimizes classification error and improves the directness and transparency of the hidden layer learning process. We focus our attention on three aspects of traditional convolutional-neural-network-type (CNN-type) architectures: (1) transparency in the effect intermediate layers have on overall classification; (2) discriminativeness and robustness of learned features, especially in early layers; (3) training effectiveness in the face of “vanishing” gradients. To combat these issues, we introduce “companion” objective functions at each hidden layer, in addition to the overall objective function at the output layer (an integrated strategy distinct from layer-wise pre-training). We also analyze our algorithm using techniques extended from stochastic gradient methods. The advantages provided by our method are evident in our experimental results, showing state-of-the-art performance on MNIST, CIFAR-10, CIFAR-100, and SVHN. Chen-Yu Lee, Saining Xie, Patrick W. Gallagher, Zhengyou Zhang, Zhuowen Tu |
AISTATS | 1 |
| 2015 | The Development of a Game-Based Formative Assessment Mathematical Algebra Tutorial App
Gwo-Haur Hwang, Chen-Yu Lee, Ting-Huan Kuo |
ICCE | 2 |
| 2014 | Region-Based Discriminative Feature Pooling for Scene Text RecognitionabstractWe present a new feature representation method for scene text recognition problem, particularly focusing on improving scene character recognition. Many existing methods rely on Histogram of Oriented Gradient (HOG) or part-based models, which do not span the feature space well for characters in natural scene images, especially given large variation in fonts with cluttered backgrounds. In this work, we propose a discriminative feature pooling method that automatically learns the most informative sub-regions of each scene character within a multi-class classification framework, whereas each sub-region seamlessly integrates a set of low-level image features through integral images. The proposed feature representation is compact, computationally efficient, and able to effectively model distinctive spatial structures of each individual character class. Extensive experiments conducted on challenging datasets (Chars74K, ICDAR'03, ICDAR'11, SVT) show that our method significantly outperforms existing methods on scene character classification and scene text recognition tasks. Chen-Yu Lee, Anurag Bhardwaj, Wei Di, Vignesh Jagadeesh, Robinson Piramuthu |
CVPR | 1 |
| 2013 | Using Augmented Reality to Assist an Interactive Multi-Language Learning System in an Elementary School
Gwo-Haur Hwang, Chen-Yu Lee, Hen-Lin Hwang, Guan-Lin Huang, Jheng-Yi Lin |
ICCE | 2 |
| 2013 | Development and Evaluation of a Problem Solving Oriented Game-Based Learning System
Hsin-Yi Liang, Song-Yu Mei, Yu-Syuan Wang, Jhih-Liang Jiang, Gwo-Haur Hwang, Chen-Yu Lee |
ICCE | 6 |
| 2012 | The Impact of Prior Knowledge and Level of Effort on the Learning Effectiveness of Subjects Using Game-Based and Traditional Certification Tutorial System
Gwo-Haur Hwang, Chen-Yu Lee, Tsung-Yen Chuang, Wei-Fang Tseng |
ICCE | 2 |
| 2011 | Ontology-Driven E-Learning System for Automated Personalized Learning ServiceabstractOntology has gained popularity in building knowledge base because the description, localization and effective reuse of software patterns and systems of patterns can be approached through an ontology-based formalism. This paper designs ontology in the mobile phone domain for the construction of knowledge base, and presents a new method for knowledge assessment in which quizzing questions are drawn using the mobile phone ontology-based knowledge base to assess the level of professional knowledge of mobile phone salespersons. Bert Chen, Chen-Yu Lee, I-Chang Tsai |
ICCE | 2 |
| 2010 | Modified Autonomous Key Management Scheme with Reduced Communication/Computation Costs in MANETabstractThe growing applications of Mobile Ad hoc Network (MANET) has made the security issue increasingly more important. B. Zhu et al. proposes a key management scheme using Shamir's secret sharing scheme to construct an Autonomous Key Management (AKM)hierarchy structure. However, Shamir's secret sharing in AKM to control key hierarchy needs larger message transmission costs. In this paper, we modify the secret sharing scheme and apply it to AKM for reducing communication and computation cost. Chu-Hsing Lin, Chen-Yu Lee |
CISIS | 2 |
| 2008 | A Learning Content Adaptation Tool with Templates for Different HandheldsabstractA large number of excellent digital learning materials have been created and distributed for e- learning. These contents are mostly designed for reading on regular PCs that have big screens, powerful computing, large storage and wide bandwidth compared to handheld devices. Therefore, a lot of studies about automatic content adaptation have been done and are proposed to overcome the drawbacks of browsing regular content with handheld devices such as pocket PCs and smartphones. But we argue that the total automatic adaptation algorithm designed by an engineer to transform Web Page presentation is still appropriate to be applied on educational content. We believe the quality of the result can not be assured and supervised by educators, and educational essence may be damaged during the real time adaptation process. This paper proposes a content adaptation tool that provides different adaptation templates to help the author automatically and efficiently reproduce high-quality learning content for specific handhelds. Furthermore, the author will not only be able to preview the adaptation result before publishing the course but also be able to adjust the template parameters manually to affect the process if they are not satisfied with the current result. Hsuan-Pu Chang, Jason C. Hung, Chun-Chia Wang, Meng-Ting Weng, Timothy K. Shih, Chen-Yu Lee |
AINA | 6 |
| 2008 | Combine Personal Blog Functionalities with LMS Using Tools Interoperability ArchitectureabstractRespecting related technologies within e-learning domain, the use of learning management system (LMS) has been considered as the most important and essential component. LMS can manage the curriculum learning content, learner's learning profile and learning process. However it lacks suitable study assistance functions and personalized interface. Accordingly LMS is hard to promote and to be utilized by learners. This paper proposed an integrated framework which combined the personal learning blog functionalities to LMS by using the tools interoperability (TI) architecture in order to develop the suitable learning functionalities and interface in LMS for learners. And we hope the TI-based blog functionalities which can be utilized by other LMS in order to improve the LMS utility rate and to prove the feasibility of TI-based blog system. Jui-Hung Chen, Timothy K. Shih, Chun-Chia Wang, Shu-Wei Yeh, Chen-Yu Lee |
AINA | 5 |
| 2008 | Erratum to "The conflict detection and resolution in knowledge merging for image annotation" [Information Processing and Management 42 (2006) 1030-1055]
Chen-Yu Lee, Von-Wun Soo |
Inf. Process. Manag. | 1 |
| 2007 | Product Outsourcing under Uncertainty: an Application of Fuzzy Real Option ApproachabstractIn recent years, in order to decrease production costs and improve competitiveness, outsourcing is adopted by many organizations. Furthermore, outsourcing is considered by more than 90 percent of investigated companies as an important part of their overall business strategy to acquire competitive advantage. Many conventional approaches take no proper account of uncertainty into the evaluation of product outsourcing, and therefore easily result in misevaluating. A new methodology, fuzzy real option approach, can be used to solve this problem. In this article, an empirical study is conducted to obtain the optimum of call price by the joint approach of fuzzy real option, weighted fuzzy real option and fuzzy decision space. The result indicates that the model developed in this study is more stable and flexible, and expected to enable the better decision making for investors. Jao-Hong Cheng, Chen-Yu Lee |
FUZZ-IEEE | 2 |
| 2006 | Using Fuzzy Analytical Hierarchy Process for Multi-criteria Evaluation Model of High-yield Bonds InvestmentabstractThe returns and risks of high-yield bond (HYB) lie between the stocks and Treasury bonds. In view of investment opportunities and the rate of return, the advantages of HYB are both lower risks and higher shares. Therefore, HYB has become one of important components in the portfolios. The purpose of this study is to find evaluation factors and their weights to aid the selection of HYB. The primary criteria to evaluate HYB are established by the literatures survey with Fuzzy Delphi Method (FDM), and then Fuzzy Analytic Hierarchy Process (FAHP) is employed to calculate the weights of these criteria, so as to build the Fuzzy Multi-criteria model of HYB investments. The results indicate a greatest weight on the dimension of economic environment, and three primary evaluation criteria are: (1) spread versus Treasuries, (2) callability, and (3) default rate. Jao-Hong Cheng, Cheng-Wei Chen, Chen-Yu Lee |
FUZZ-IEEE | 3 |
| 2004 | An Image Annotation Guide Agent
Chen-Yu Lee, Von-Wun Soo, Yi-Ting Fu |
PRIMA | 1 |
| 2001 | Gaz-Guide: Agent-Mediated Information Retrieval for Official Gazettes
Jyi-Shane Liu, Von-Wun Soo, Chia-Ning Chiang, Chen-Yu Lee |
PRIMA | 4 |
| 1995 | An O(log n) Parallel Algorithm for Constructing a Spanning Tree on Permutation Graphs
Yue-Li Wang, Hon-Chan Chen, Chen-Yu Lee |
Inf. Process. Lett. | 3 |