EDBT 2026 Demo / reviewers in the wild / expert
Licheng Yu
dblp:32/10805
· DBLP profile ↗
55ranked-venue papers
13as first author
25since 2021 · last 2026
0000-0002-4943-6732ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 7 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 18 since 2021Systems, architecture and hardware · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction FollowingabstractYun He, Wenzhe Li, Hejia Zhang, Songlin Li, Karishma Mandyam, Sopan Khosla, Yuanhao Xiong, Nanshu Wang, Xiaoliang Peng, Beibin Li, Shengjie Bi, Shishir G Patil, Qi Qi, Shengyu Feng, Julian Katz-Samuels, Richard Yuanzhe Pang, Sujan Kumar Gonugondla, Hunter Lang, Yue Yu, Yundi Qian, Maryam Fazel-Zarandi, Licheng Yu, Amine Benhalloum, Hany Hassan Awadalla, Manaal Faruqui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Karishma Mandyam, Sopan Khosla, Yuanhao Xiong, Nanshu Wang, Xiaoliang Peng, Beibin Li, Shengjie Bi, Shishir G. Patil, Shengyu Feng, Julian Katz-Samuels, Richard Yuanzhe Pang, Sujan K. Gonugondla, Hunter Lang, Yue Yu 0009, Yundi Qian, Maryam Fazel-Zarandi, Licheng Yu, Amine Benhalloum, Hany Hassan, Manaal Faruqui |
ACL (1) | 22 |
| 2026 | KScaNN: Scalable Approximate Nearest Neighbor Search on KunpengabstractApproximate Nearest Neighbor Search (ANNS) is a cornerstone algorithm for information retrieval, recommendation systems, and machine learning applications. While x86-based architectures have historically dominated this domain, the increasing adoption of ARM-based servers in industry presents a critical need for ANNS solutions optimized on ARM architectures. A naive port of existing x86 ANNS algorithms to ARM platforms results in a substantial performance deficit, failing to leverage the unique capabilities of the underlying hardware. To address this challenge, we introduce KScaNN, a novel ANNS algorithm co-designed for the Kunpeng 920 ARM architecture. KScaNN embodies a holistic approach that synergizes sophisticated, data aware algorithmic refinements with carefully-designed hardware specific optimizations. Its core contributions include: 1) novel algorithmic techniques, including a hybrid intra-cluster search strategy and an improved PQ residual calculation method, which optimize the search process at a higher level; 2) an ML-driven adaptive search module that provides adaptive, per-query tuning of search parameters, eliminating the inefficiencies of static configurations; and 3) highly-optimized SIMD kernels for ARM that maximize hardware utilization for the critical distance computation workloads. The experimental results demonstrate that KScaNN not only closes the performance gap but establishes a new standard, achieving up to a 1.63x speedup over the fastest x86-based solution. This work provides a definitive blueprint for achieving leadership-class performance for vector search on modern ARM architectures and underscores Oleg Senkevich, Siyang Xu, Tianyi Jiang, Alexander Radionov, Jan Tabaszewski, Dmitriy Malyshev, Daihao Xue, Licheng Yu, Weidi Zeng, Xin Yao 0008, Siyu Huang, Gleb Neshchetkin, Qiuling Pan, Yaoyao Fu |
ICDE | 9 |
| 2025 | ROICtrl: Boosting Instance Control for Visual GenerationabstractNatural language often struggles to accurately associate positional and attribute information with multiple instances, which limits current text-based visual generation models to simpler compositions featuring only a few dominant instances. To address this limitation, this work enhances diffusion models by introducing regional instance control, where each instance is governed by a bounding box paired with a free-form caption. Previous methods in this area typically rely on implicit position encoding or explicit attention masks to separate regions of interest (ROIs), resulting in either inaccurate coordinate injection or large computational overhead. Inspired by ROI-Align in object detection, we introduce a complementary operation called ROI-Unpool. Together, ROI-Align and ROI- Unpool enable explicit, efficient, and accurate ROI manipulation on high-resolution feature maps for visual generation. Building on ROI-Unpool, we propose ROICtrl, an adapter for pretrained diffusion models that enables precise regional instance control. ROICtrl is compatible with community-finetuned diffusion models, as well as with existing spatial-based add-ons (e.g., ControlNet, T2I- Adapter) and embedding-based add-ons (e.g., IP-Adapter, ED-LoRA), extending their applications to multi-instance generation. Experiments show that ROICtrl achieves superior performance in regional instance control while significantly reducing computational costs. Yuchao Gu, Yipin Zhou, Yunfan Ye, Yixin Nie, Licheng Yu, Pingchuan Ma 0002, Qinghong Lin, Zheng Shou 0001 |
CVPR | 5 |
| 2025 | Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMsabstractLong-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited context windows. In this work, we introduce VideoMindPalace, a new framework inspired by the "Mind Palace", which organizes critical video moments into a topologically structured semantic graph. VideoMindPalace organizes key information through (i) hand-object tracking and interaction, (ii) clustered activity zones representing specific areas of recurring activities, and (iii) environment layout mapping, allowing natural language parsing by LLMs to provide grounded insights on spatio-temporal and 3D context. In addition, we propose the Video Mind-Palace Benchmark (VMB), to assess human-like reasoning, including spatial localization, temporal reasoning, and layout-aware sequential understanding. Evaluated on VMB and established video QA datasets, including EgoSchema, NExT-QA, IntentQA, and the Active Memories Benchmark, VideoMindPalace demonstrates notable gains in spatiotemporal coherence and human-aligned reasoning, advancing long-form video analysis capabilities in VLMs. Project website: https://oodbag.github.io/VMP/. Zeyi Huang, Yuyang Ji, Nikhil Mehta 0002, Tong Xiao 0003, Donghyun Lee 0004, Sigmund Vanvalkenburgh, Shengxin Zha, Bolin Lai, Licheng Yu, Yong Jae Lee, Miao Liu 0007 |
CVPR | 10 |
| 2025 | Accelerating Multimodal Large Language Models by Searching Optimal Vision Token ReductionabstractPrevailing Multimodal Large Language Models (MLLMs) encode the input image(s) as vision tokens and feed them into the language backbone, similar to how Large Language Models (LLMs) process the text tokens. However, the number of vision tokens increases quadratically as the image resolutions, leading to huge computational costs. In this paper, we consider improving MLLM’s efficiency from two scenarios, (I) Reducing computational cost without degrading the performance. (II) Improving the performance with given budgets. We start with our main finding that the ranking of each vision token sorted by attention scores is similar in each layer except the first layer. Based on it, we assume that the number of essential top vision tokens does not increase along layers. Accordingly, for Scenario I, we propose a greedy search algorithm (G-Search) to find the least number of vision tokens to keep at each layer from the shallow to the deep. Interestingly, G-Search is able to reach the optimal reduction strategy based on our assumption. For Scenario II, based on the reduction strategy from G-Search, we design a parametric sigmoid function (P-Sigmoid) to guide the reduction at each layer of the MLLM, whose parameters are optimized by Bayesian Optimization. Extensive experiments demonstrate that our approach can significantly accelerate those popular MLLMs, e.g. LLaVA, and InternVL2 models, by more than 2⇥ without performance drops. Our approach also far outperforms other token reduction methods when budgets are limited, achieving a better trade-off between efficiency and effectiveness. Shiyu Zhao 0001, Zhenting Wang, Felix Juefei-Xu, Xide Xia, Miao Liu 0007, Mingfu Liang, Dimitris N. Metaxas, Licheng Yu |
CVPR | 10 |
| 2025 | Apollo: An Exploration of Video Understanding in Large Multimodal ModelsabstractDespite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), what drives their video perception remains poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models and limited open research hinder the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs. We begin by critically examining the primary contributors to the high computational requirements associated with video-LMM research and discover Scaling Consistency, wherein design and training decisions made on smaller models and datasets (up to a critical size) effectively transfer to larger models. Leveraging these insights, we explored many video-specific aspects of video-LMMs, including video sampling, architectures, data composition, training schedules, and more. Guided by these findings, we introduce Apollo, a state-of-the-art family of LMMs that achieve superior performance across different model sizes. Our models process over 1-hour videos efficiently, with the 3B parameter variant outperforming most existing 7B models. Apollo-7B is state-of-the-art compared to 7B LMMs with a 70.9 on MLVU, and 63.3 on Video-MME. Orr Zohar, Yann Dubois, Nikhil Mehta 0002, Tong Xiao 0003, Philippe Hansen-Estruch, Licheng Yu, Felix Juefei-Xu, Serena Yeung-Levy, Xide Xia |
CVPR | 7 |
| 2024 | VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point CorrespondenceabstractCurrent diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignment. However, these approaches are often in-effective when the target edit involves a shape change. To embark on video editing with shape change, we explore customized video subject swapping in this work, where we aim to replace the main subject in a source video with a target subject having a distinct identity and potentially different shape. In contrast to previous methods that rely on dense correspondences, we introduce the Video Swap framework that exploits semantic point correspondences, inspired by our observation that only a small number of semantic points are necessary to align the subject's motion trajectory and modify its shape. We also introduce various user-point interactions (e.g., removing points and dragging points) to address various semantic point correspondence. Extensive experiments demonstrate state-of-the-art video subject swapping results across a variety of real-world videos. Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao 0001, Jay Zhangjie Wu, Junhao Zhang 0001, Zheng Shou 0001, Kevin Tang |
CVPR | 4 |
| 2024 | FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisabstractDiffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. How-ever, the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining temporal consistency across video frames. This paper proposes a consistent V2V synthesis framework by jointly leveraging spatial conditions and temporal optical flow clues within the source video. Contrary to prior methods that strictly adhere to optical flow, our approach harnesses its benefits while handling the imperfection in flow estimation. We encode the optical flow via warping from the first frame and serve it as a supplementary reference in the diffusion model. This enables our model for video synthesis by editing the first frame with any prevalent I2I models and then propagating edits to successive frames. Our V2V model, Flow Vid, demon-strates remarkable properties: (1) Flexibility: Flow Vid works seamlessly with existing I2I models, facilitating various modifications, including stylization, object swaps, and local edits. (2) Efficiency: Generation of a 4-second video with 30 FPS and 512×512 resolution takes only 1.5 minutes, which is 3.1×, 7.2×, and 10.5× faster than CoDeF, Rerender, and TokenFlow, respectively. (3) High-quality: In user studies, our FlowVid is preferred 45.7% of the time, outperforming CoDeF (3.5%), Rerender (10.2%), and TokenFlow (40.4%). Bichen Wu, Jialiang Wang 0001, Licheng Yu, Ishan Misra, Jia-Bin Huang 0001, Peizhao Zhang, Peter Vajda, Diana Marculescu |
CVPR | 4 |
| 2024 | Fairy: Fast Parallelized Instruction-Guided Video-to-Video SynthesisabstractIn this paper, we introduce Fairy, a minimalist yet ro-bust adaptation of image-editing diffusion models, enhancing them for video editing applications. Our approach centers on the concept of anchor-based cross-frame attention, a mechanism that implicitly propagates diffusion features across frames, ensuring superior temporal coherence and high-fidelity synthesis. Fairy not only addresses limitations of previous models on memory and processing speed, but also improves temporal consistency through a unique data augmentation strategy. This strategy renders the model equivariant to affine transformations in both source and target images. Remarkably efficient, Fairy generates 120- frame 512×384 videos (4-second duration at 30 FPS) in just 14 seconds, outpacing prior works by at least 44×. A comprehensive user study, involving 1000 generated samples, confirms that our approach delivers superior quality, decisively outperforming established methods. Bichen Wu, Ching-Yao Chuang, Kapil Krishnakumar, Tong Xiao 0003, Licheng Yu, Peter Vajda |
CVPR | 8 |
| 2024 | AVID: Any-Length Video Inpainting with Diffusion ModelabstractRecent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain, there have been fewer works regarding textguided video inpainting. Given a video, a masked region at its initial frame, and an editing prompt, it requires a model to do infilling at each frame following the editing guidance while keeping the out-of-mask region intact. There are three main challenges in text-guided video inpainting: (i) temporal consistency of the edited video, (ii) supporting different inpainting types at different structural fidelity levels, and (iii) dealing with variable video length. To address these challenges, we introduce Any-Length Video Inpainting with Diffusion Model, dubbed as AVID. At its core, our model is equipped with effective motion modules and adjustable structure guidance, for fixed-length video inpainting. Building on top of that, we propose a novel Temporal MultiDiffusion sampling pipeline with a middle-frame attention guidance mechanism, facilitating the generation of videos with any desired duration. Our comprehensive experiments show our model can robustly deal with various inpainting types at different video duration ranges, with high quality11More visualization results are made publicly available here. Bichen Wu, Yaqiao Luo, Luxin Zhang, Peter Vajda, Dimitris N. Metaxas, Licheng Yu |
CVPR | 9 |
| 2024 | Layout-Agnostic Scene Text Image Synthesis with Diffusion ModelsabstractWhile diffusion models have significantly advanced the quality of image generation, their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene text generation are typically limited by their reliance on an intermediate layout output. This dependency often results in a constrained diversity of text styles and fonts, an inherent limitation stemming from the deterministic nature of the layout generation phase. To address these challenges, this paper introduces Scene TextGen, a novel diffusion-based model specifically designed to circumvent the need for a predefined layout stage. By doing so, Scene-TextGen facilitates a more natural and varied representation of text. The novelty of SceneTextGen lies in its integration of three key components: a character-level encoder for capturing detailed typographic properties, coupled with a character-level instance segmentation model and a word-level spotting model to address the issues of unwanted text generation and minor character inaccuracies. We validate the performance of our method by demonstrating improved character recognition rates on generated images across different public visual text datasets in comparison to both standard diffusion based methods and text specific methods. Qilong Zhangli, Jindong Jiang, Di Liu 0003, Licheng Yu, Xiaoliang Dai, Ankit Ramchandani, Guan Pang, Dimitris N. Metaxas, Praveen Krishnan |
CVPR | 4 |
| 2024 | Ameli: Enhancing Multimodal Entity Linking with Fine-Grained AttributesabstractBarry Yao, Sijia Wang, Yu Chen, Qifan Wang, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Barry Menglong Yao, Yu Chen 0022, Qifan Wang 0001, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang |
EACL (1) | 7 |
| 2024 | Text-to-Sticker: Style Tailoring Latent Diffusion Models for Human Expression
Animesh Sinha, Anmol Kalia, Arantxa Casanova, Elliot Blanchard, David Yan, Winnie Zhang, Tony Nelli, Hardik Shah, Licheng Yu, Mitesh Kumar Singh, Ankit Ramchandani, Maziar Sanjabi, Sonal Gupta, Amy Bearman, Dhruv Mahajan 0001 |
ECCV (70) | 11 |
| 2023 | Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video GenerationabstractGenerating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head and tail is also crucial, but they have rarely been explored for video completion. Since there could be different outcomes from the hints of just a few frames, a system that can follow natural language to perform video completion may significantly improve controllability. Inspired by this, we introduce a novel task, text-guided video completion (TVC), which requests the model to generate a video from partial frames guided by an instruction. We then propose Multimodal Masked Video Generation (MMVG) to address this TVC task. During training, MMVG discretizes the video frames into visual tokens and masks most of them to perform video completion from any time point. At inference time, a single MMVG model can address all 3 cases of TVC, including video prediction, rewind, and infilling, by applying corresponding masking conditions. We evaluate MMVG in various video scenarios, including egocentric, animation, and gaming. Extensive experimental results indicate that MMVG is effective in generating high-quality visual appearances with text guidance for TVC. Tsu-Jui Fu, Licheng Yu, Ning Zhang 0014, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, Sean Bell |
CVPR | 2 |
| 2023 | FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion TasksabstractIn the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in each individual input/output format and dataset size. It has been common to design a task-specific model and fine-tune it independently from a pre-trained V+l model (e.g., CLIP). This results in parameter inefficiency and inability to exploit inter-task relatedness. To address such issues, we propose a novel FAshion-focused Multi-task Efficient learning method for Vision-and-Language tasks (FAME-ViL) in this work. Compared with existing approaches, FAME-ViL applies a single model for multiple heterogeneous fashion tasks, therefore being much more parameter-efficient. It is enabled by two novel components: (1) a task-versatile architecture with cross-attention adapters and task-specific adapters integrated into a unified V+L model, and (2) a stable and effective multi-task training strategy that supports learning from heterogeneous data and prevents negative transfer. Extensive experiments on four fashion tasks show that our FAME-ViL can save 61.5% of parameters over alternatives, while significantly outperforming the conventional independently trained single-task models. Code is available at https://github.com/BrandonHanx/FAME-ViL. Xiatian Zhu, Licheng Yu, Li Zhang 0040, Yi-Zhe Song, Tao Xiang 0002 |
CVPR | 3 |
| 2023 | Learning Procedure-aware Video Representation from Instructional Videos and Their NarrationsabstractThe abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn video representation that encodes both action steps and their temporal ordering, based on a large-scale dataset of web instructional videos and their narrations, without using human annotations. Our method jointly learns a video representation to encode individual step concepts, and a deep probabilistic model to capture both temporal dependencies and immense individual variations in the step ordering. We empirically demonstrate that learning temporal ordering not only enables new capabilities for procedure reasoning, but also reinforces the recognition of individual steps. Our model significantly advances the state-of-the-art results on step classification (+2.8%/+3.3% on COIN / EPIC-Kitchens) and step forecasting (+7.4% on COIN). Moreover, our model attains promising results in zero-shot inference for step classification and fore-casting, as well as in predicting diverse and plausible steps for incomplete procedures. Our code is available at https://github.com/facebookresearch/ProcedureVRL. Yiwu Zhong, Licheng Yu, Shangwen Li, Xueting Yan, Yin Li 0003 |
CVPR | 2 |
| 2023 | CiT: Curation in Training for Effective Vision-Language DataabstractLarge vision-language models are generally applicable to many downstream tasks, but come at an exorbitant training cost that only large institutions can afford. This paper trades generality for efficiency and presents Curation in Training (CiT), a simple and efficient vision-text learning algorithm that couples a data objective into training. CiT automatically yields quality data to speed-up contrastive image-text training and alleviates the need for an offline data filtering pipeline, allowing broad data sources (including raw image-text pairs from the web). CiT contains two loops: an outer loop curating the training data and an inner loop consuming the curated training data. The text encoder connects the two loops. Given metadata for tasks of interest, e.g., class names, and a large pool of image-text pairs, CiT alternatively selects relevant training data from the pool by measuring the similarity of their text embeddings and embeddings of the metadata. In our experiments, we observe that CiT can speed up training by over an order of magnitude, especially if the raw data size is large. Hu Xu 0001, Saining Xie, Po-Yao Huang 0001, Licheng Yu, Russell Howes, Gargi Ghosh, Luke Zettlemoyer, Christoph Feichtenhofer |
ICCV | 4 |
| 2023 | RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data
Sangwoo Mo, Jong-Chyi Su, Chih-Yao Ma, Mido Assran, Ishan Misra, Licheng Yu, Sean Bell |
ICLR | 6 |
| 2022 | Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentabstractVision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel imagetext data, which is costly to collect, compared to image-only or text-only data. In this paper, we explore unsupervised Vision-and-Language pre-training (UVLP) to learn the cross-modal representation from non-parallel image and text datasets. We found two key factors that lead to good unsupervised V + L pre-training without parallel data: (i) joint image-and-text input (ii) overall imagetext alignment (even for non-parallel data). Accordingly, we propose a novel unsupervised V + L pre-training curriculum for non-parallel texts and images. We first construct a weakly aligned imagetext corpus via a retrieval-based approach, then apply a set of multi-granular alignment pre-training tasks, including region-to-tag, region-to-phrase, and image-to-sentence alignment, to bridge the gap between the two modalities. A comprehensive ablation study shows each granularity is helpful to learn a stronger pre-trained model. We adapt our pre-trained model to a set of V+L downstream tasks, including VQA, NLVR2, Visual Entailment, and Ref-COCO+. Our model achieves the state-of-art performance in all these tasks under the unsupervised setting. Mingyang Zhou 0004, Licheng Yu, Amanpreet Singh, Mengjiao Wang 0002, Zhou Yu 0005 |
CVPR | 2 |
| 2022 | FashionViL: Fashion-Focused Vision-and-Language Representation Learning
Licheng Yu, Xiatian Zhu, Li Zhang 0040, Yi-Zhe Song, Tao Xiang 0002 |
ECCV (35) | 2 |
| 2022 | GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
Difei Gao, Licheng Yu, Weixian Lei, Matt Feiszli, Zheng Shou 0001 |
ECCV (35) | 3 |
| 2022 | FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningabstractMultimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems-e.g., retrieving a fashion item given a reference image plus text feedback from a user.Prior works on multimodal fashion tasks have either been limited by the data in individual benchmarks, or have leveraged generic vision-and-language pre-training but have not taken advantage of the characteristics of fashion data.Additionally, these works have mainly been restricted to multimodal understanding tasks.To address these gaps, we make two key contributions.First, we propose a novel fashion-specific pre-training framework based on weakly-supervised triplets constructed from fashion image-text pairs.We show the tripletbased tasks are an effective addition to standard multimodal pre-training tasks.Second, we propose a flexible decoder-based model architecture capable of both fashion retrieval and captioning tasks.Together, our model design and pre-training approach are competitive on a diverse set of fashion tasks, including crossmodal retrieval, image retrieval with text feedback, image captioning, relative image captioning, and multimodal categorization. Suvir Mirchandani, Licheng Yu, Mengjiao Wang 0002, Animesh Sinha, Wenwen Jiang, Tao Xiang 0002 |
EMNLP | 2 |
| 2022 | CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni RetrievalabstractWe introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, image+text), and having the capability to generalize to a wide range of tasks, including Multimodal Categorization, Image-Text Retrieval, Query-to-Product Retrieval, Image-to-Product Retrieval, etc. We follow the pre-training + fine-tuning training regime and present 5 effective pre-training tasks on image-text pairs. To embrace more common and diverse commerce data with text-to-multimodal, image-to-multimodal, and multimodal-to-multimodal mapping, we propose another 9 novel cross-modal and cross-pair retrieval tasks, called Omni-Retrieval pre-training. We also propose a novel approach of modality randomization to dynamically adjust our model under different efficiency constraints. The pre-training is conducted in an efficient manner with only two forward/backward updates for the combined 14 tasks. Extensive experiments and analysis show the effectiveness of each task. When combining all pre-training tasks, our model achieves state-of-the-art performance on 7 commerce-related downstream tasks after fine-tuning. Licheng Yu, Animesh Sinha, Mengjiao Wang 0002, Tamara L. Berg |
KDD | 1 |
| 2021 | Connecting What To Say With Where To Look by Modeling Human Attention TracesabstractWe introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation frame-work [31], where each word of a given caption is paired with a mouse trace segment. We propose two novel tasks: (1) predict a trace given an image and caption (i.e., visual grounding), and (2) predict a caption and a trace given only an image. Learning the grounding of each word is challenging, due to noise in the human-provided traces and the presence of words that cannot be meaningfully visually grounded. We present a novel model architecture that is jointly trained on dual tasks (controlled trace generation and controlled caption generation). To evaluate the quality of the generated traces, we propose a local bipartite matching (LBM) distance metric which allows the comparison of two traces of different lengths. Extensive experiments show our model is robust to the imperfect training data and outperforms the baselines by a clear margin. More-over, we demonstrate that our model pre-trained on the pro-posed tasks can be also beneficial to the downstream task of COCO’s guided image captioning. Our code1and project page2are publicly available. Zihang Meng, Licheng Yu, Tamara L. Berg, Babak Damavandi, Amy Bearman |
CVPR | 2 |
| 2021 | Assistive supernumerary grasping with the back of the handabstractThe Dorsal Grasper, an assistive wearable grasping device, incorporates supernumerary fingers and an artificial palm with the forearm and back of the hand, respectively. It enables power wrap grasping and adduction pinching with its V-shaped soft fingers. Designed with C6/C7 spinal cord injury in mind, it takes advantage of active wrist extension that remains in this population after injury. We propose that allowing the operator to actively participate in applying grasp forces on the object, using the back of the hand, enables intuitive, fast and reliable grasping relevant for the execution of activities of daily living. Functional grasping is tested in three normative subjects and a person with C6 SCI using the Grasp and Release Test. Results indicate that this device provides promising performance on a subset of objects that complements the existing compensatory strategies used by people with C6/C7 SCI. We find that the addition of the artificial palm is important for increasing maximum grip strength, by increasing contact friction and protecting the opisthenar. Jungpyo Lee, Licheng Yu, Lucie Derbier, Hannah Stuart |
ICRA | 2 |
| 2020 | TVQA+: Spatio-Temporal Grounding for Video Question AnsweringabstractWe present the task of Spatio-Temporal Video Question Answering, which requires intelligent systems to simultaneously retrieve relevant moments and detect referenced visual concepts (people and objects) to answer natural language questions about videos. We first augment the TVQA dataset with 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers. We name this augmented version as TVQA+. We then propose Spatio-Temporal Answerer with Grounded Evidence (STAGE), a unified framework that grounds evidence in both spatial and temporal domains to answer questions about videos. Comprehensive experiments and analyses demonstrate the effectiveness of our framework and how the rich annotations in our TVQA+ dataset can contribute to the question answering task. Moreover, by performing this joint task, our model is able to produce insightful and interpretable spatio-temporal attention visualizations. Jie Lei 0003, Licheng Yu, Tamara L. Berg, Mohit Bansal |
ACL | 2 |
| 2020 | BachGAN: High-Resolution Image Synthesis From Salient Object LayoutabstractWe propose a new task towards more practical applications for image generation - high-quality image synthesis from salient object layout. This new setting requires users to provide only the layout of salient objects (i.e., foreground bounding boxes and categories) and lets the model complete the drawing with an invented background and a matching foreground. Two main challenges spring from this new task: (i) how to generate fine-grained details and realistic textures without segmentation map input; and (ii) how to create and weave a background into standalone objects in a seamless way. To tackle this, we propose Background Hallucination Generative Adversarial Network (BachGAN), which leverages a background retrieval module to first select a set of segmentation maps from a large candidate pool, then encodes these candidate layouts via a background fusion module to hallucinate a suitable background for the given objects. By generating the hallucinated background representation dynamically, our model can synthesize high-resolution images with both photo-realistic foreground and integral background. Experiments on Cityscapes and ADE20K datasets demonstrate the advantage of BachGAN over existing approaches, measured on both visual fidelity of generated images and visual alignment between output images and input layouts. Yandong Li, Yu Cheng 0001, Zhe Gan, Licheng Yu, Liqiang Wang 0001, Jingjing Liu 0001 |
CVPR | 4 |
| 2020 | Violin: A Large-Scale Dataset for Video-and-Language InferenceabstractWe introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the video content, a model needs to infer whether the hypothesis is entailed or contradicted by the given video clip. A new large-scale dataset, named Violin (VIdeO-and-Language INference), is introduced for this task, which consists of 95,322 video-hypothesis pairs from 15,887 video clips, spanning over 582 hours of video. These video clips contain rich content with diverse temporal dynamics, event shifts, and people interactions, collected from two sources: (i) popular TV shows, and (ii) movie clips from YouTube channels. In order to address our new multimodal inference task, a model is required to possess sophisticated reasoning skills, from surface-level grounding (e.g., identifying objects and characters in the video) to in-depth commonsense reasoning (e.g., inferring causal relations of events in the video). We present a detailed analysis of the dataset and an extensive evaluation over many strong baselines, providing valuable insights on the challenges of this new task. Jingzhou Liu, Wenhu Chen, Yu Cheng 0001, Zhe Gan, Licheng Yu, Yiming Yang 0002, Jingjing Liu 0001 |
CVPR | 5 |
| 2020 | Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Jize Cao, Zhe Gan, Yu Cheng 0001, Licheng Yu, Yen-Chun Chen 0001, Jingjing Liu 0001 |
ECCV (6) | 4 |
| 2020 | UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen 0001, Licheng Yu, Ahmed El Kholy, Faisal Ahmed 0001, Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001 |
ECCV (30) | 3 |
| 2020 | TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
Jie Lei 0003, Licheng Yu, Tamara L. Berg, Mohit Bansal |
ECCV (21) | 2 |
| 2020 | What is More Likely to Happen Next? Video-and-Language Future Event PredictionabstractGiven a video with aligned dialogue, people can often infer what is more likely to happen next.Making such predictions requires not only a deep understanding of the rich dynamics underlying the video and dialogue, but also a significant amount of commonsense knowledge.In this work, we explore whether AI models are able to learn to make such multimodal commonsense nextevent predictions.To support research in this direction, we collect a new dataset, named Video-and-Language Event Prediction (VLEP), with 28,726 future event prediction examples (along with their rationales) from 10,234 diverse TV Show and YouTube Lifestyle Vlog video clips.In order to promote the collection of non-trivial challenging examples, we employ an adversarial humanand-model-in-the-loop data collection procedure.We also present a strong baseline incorporating information from video, dialogue, and commonsense knowledge.Experiments show that each type of information is useful for this challenging task, and that compared to the high human performance on VLEP, our model provides a good starting point but leaves large room for future work. 1 Jie Lei 0003, Licheng Yu, Tamara L. Berg, Mohit Bansal |
EMNLP (1) | 2 |
| 2020 | HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingabstractWe present HERO, a novel framework for large-scale video+language omnirepresentation learning.HERO encodes multimodal inputs in a hierarchical structure, where local context of a video frame is captured by a Cross-modal Transformer via multimodal fusion, and global video context is captured by a Temporal Transformer.In addition to standard Masked Language Modeling (MLM) and Masked Frame Modeling (MFM) objectives, we design two new pre-training tasks: (i) Video-Subtitle Matching (VSM), where the model predicts both global and local temporal alignment; and (ii) Frame Order Modeling (FOM), where the model predicts the right order of shuffled video frames.HERO is jointly trained on HowTo100M and large-scale TV datasets to gain deep understanding of complex social dynamics with multi-character interactions.Comprehensive experiments demonstrate that HERO achieves new state of the art on multiple benchmarks over Text-based Video/Video-moment Retrieval, Video Question Answering (QA), Video-and-language Inference and Video Captioning tasks across different domains.We also introduce two new challenging benchmarks How2QA and How2R for Video QA and Retrieval, collected from diverse video content over multimodalities. 1 Yen-Chun Chen 0001, Yu Cheng 0001, Zhe Gan, Licheng Yu, Jingjing Liu 0001 |
EMNLP (1) | 5 |
| 2019 | Multi-Target Embodied Question AnsweringabstractEmbodied Question Answering (EQA) is a relatively new task where an agent is asked to answer questions about its environment from egocentric perception. EQA as introduced in [8] makes the fundamental assumption that every question, e.g., ``what color is the car?", has exactly one target (``car") being inquired about. This assumption puts a direct limitation on the abilities of the agent. We present a generalization of EQA -- Multi-Target EQA (MT-EQA). Specifically, we study questions that have multiple targets in them, such as ``Is the dresser in the bedroom bigger than the oven in the kitchen?", where the agent has to navigate to multiple locations (``dresser in bedroom", ``oven in kitchen") and perform comparative reasoning (``dresser" bigger than ``oven") before it can answer a question. Such questions require the development of entirely new modules or components in the agent. To address this, we propose a modular architecture composed of a program generator, a controller, a navigator, and a VQA module. The program generator converts the given question into sequential executable sub-programs; the navigator guides the agent to multiple locations pertinent to the navigation-related sub-programs; and the controller learns to select relevant observations along its path. These observations are then fed to the VQA module to predict the answer. We perform detailed analysis for each of the model components and show that our joint model can outperform previous methods and strong baselines by a significant margin. Licheng Yu, Xinlei Chen, Georgia Gkioxari, Mohit Bansal, Tamara L. Berg, Dhruv Batra |
CVPR | 1 |
| 2018 | MAttNet: Modular Attention Network for Referring Expression ComprehensionabstractIn this paper, we address referring expression comprehension: localizing an image region described by a natural language expression. While most recent work treats expressions as a single unit, we propose to decompose them into three modular components related to subject appearance, location, and relationship to other objects. This allows us to flexibly adapt to expressions containing different types of information in an end-to-end framework. In our model, which we call the Modular Attention Network (MAttNet), two types of attention are utilized: language-based attention that learns the module weights as well as the word/phrase attention that each module should focus on; and visual attention that allows the subject and relationship modules to focus on relevant image components. Module weights combine scores from all three modules dynamically to output an overall score. Experiments show that MAttNet outperforms previous state-of-the-art methods by a large margin on both bounding-box-level and pixel-level comprehension tasks. Demo1 and code2 are provided. Licheng Yu, Zhe Lin 0001, Xiaohui Shen, Jimei Yang, Xin Lu 0006, Mohit Bansal, Tamara L. Berg |
CVPR | 1 |
| 2018 | TVQA: Localized, Compositional Video Question AnsweringabstractRecent years have witnessed an increasing interest in image-based question-answering (QA) tasks.However, due to data limitations, there has been much less work on video-based QA.In this paper, we present TVQA, a largescale video QA dataset based on 6 popular TV shows.TVQA consists of 152,545 QA pairs from 21,793 clips, spanning over 460 hours of video.Questions are designed to be compositional in nature, requiring systems to jointly localize relevant moments within a clip, comprehend subtitle-based dialogue, and recognize relevant visual concepts.We provide analyses of this new dataset as well as several baselines and a multi-stream end-to-end trainable neural network framework for the TVQA task.The dataset is publicly available at http://tvqa.cs.unc.edu. Jie Lei 0003, Licheng Yu, Mohit Bansal, Tamara L. Berg |
EMNLP | 2 |
| 2018 | Last level cache layout remapping for heterogeneous systems
Licheng Yu, Tianzhou Chen, Minghui Wu 0001, Xueqing Lou |
J. Syst. Archit. | 1 |
| 2018 | From image to language and back againabstractWork in computer vision and natural language processing involving images and text has been experiencing explosive growth over the past decade, with a particular boost coming from the neural network revolution. The present volume brings together five research articles from several different corners of the area: multilingual multimodal image description (Franket al.), multimodal machine translation (Madhyasthaet al., Franket al.), image caption generation (Madhyasthaet al., Tantiet al.), visual scene understanding (Silbereret al.), and multimodal learning of high-level attributes (Sorodocet al.). In this article, we touch upon all of these topics as we review work involving images and text under the three main headings of image description (Section 2), visually grounded referring expression generation (REG) and comprehension (Section 3), and visual question answering (VQA) (Section 4). Anya Belz, Tamara L. Berg, Licheng Yu |
Nat. Lang. Eng. | 3 |
| 2018 | Physics-Inspired Garment Recovery from a Single-View ImageabstractMost recent garment capturing techniques rely on acquiring multiple views of clothing, which may not always be readily available, especially in the case of pre-existing photographs from the web. As an alternative, we propose a method that is able to compute a 3D model of a human body and its outfit from a single photograph with little human interaction. Our algorithm is not only able to capture the global shape and overall geometry of the clothing, it can also extract the physical properties (i.e., material parameters needed for simulation) of cloth. Unlike previous methods using full 3D information (i.e., depth, multi-view images, or sampled 3D geometry), our approach achieves garment recovery from a single-view image by using physical, statistical, and geometric priors and a combination of parameter estimation, semantic parsing, shape/pose recovery, and physics-based cloth simulation. We demonstrate the effectiveness of our algorithm by re-purposing the reconstructed garments for virtual try-on and garment transfer applications and for cloth animation on digital characters. Zherong Pan, Tanya Amert, Ke Wang 0021, Licheng Yu, Tamara L. Berg, Ming C. Lin |
ACM Trans. Graph. | 5 |
| 2017 | A Joint Speaker-Listener-Reinforcer Model for Referring ExpressionsabstractReferring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generation. Our model is composed of three modules: speaker, listener, and reinforcer. The speaker generates referring expressions, the listener comprehends referring expressions, and the reinforcer introduces a reward function to guide sampling of more discriminative expressions. The listener-speaker modules are trained jointly in an end-to-end learning framework, allowing the modules to be aware of one another during learning while also benefiting from the discriminative reinforcer's feedback. We demonstrate that this unified framework and training achieves state-of-the-art results for both comprehension and generation on three referring expression datasets. Licheng Yu, Hao Tan 0002, Mohit Bansal, Tamara L. Berg |
CVPR | 1 |
| 2017 | Hierarchically-Attentive RNN for Album Summarization and StorytellingabstractWe address the problem of end-to-end visual storytelling.Given a photo album, our model first selects the most representative (summary) photos, and then composes a natural language story for the album.For this task, we make use of the Visual Storytelling dataset and a model composed of three hierarchically-attentive Recurrent Neural Nets (RNNs) to: encode the album photos, select representative (summary) photos, and compose the story.Automatic and human evaluations show our model achieves better performance on selection, generation, and retrieval than baselines. Licheng Yu, Mohit Bansal, Tamara L. Berg |
EMNLP | 1 |
| 2017 | Enable back memory and global synchronization on LLC buffer
Licheng Yu, Yulong Pei, Tianzhou Chen, Xueqing Lou, Minghui Wu 0001, Tiefei Zhang |
J. Supercomput. | 1 |
| 2016 | Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Alexander C. Berg, Tamara L. Berg |
ECCV (2) | 1 |
| 2016 | Architecture supported register stash for GPGPU
Licheng Yu, Yulong Pei, Tianzhou Chen, Minghui Wu 0001 |
J. Parallel Distributed Comput. | 1 |
| 2015 | Dictionary Learning with Mutually Reinforcing Group-Graph StructuresabstractIn this paper, we propose a novel dictionary learning method in the semi-supervised setting by dynamically coupling graph and group structures. To this end, samples are represented by sparse codes inheriting their graph structure while the labeled samples within the same class are represented with group sparsity, sharing the same atoms of the dictionary. Instead of statically combining graph and group structures, we take advantage of them in a mutually reinforcing way — in the dictionary learning phase, we introduce the unlabeled samples into groups by an entropy-based method and then update the corresponding local graph, resulting in a more structured and discriminative dictionary. We analyze the relationship between the two structures and prove the convergence of our proposed method. Focusing on image classification task, we evaluate our approach on several datasets and obtain superior performance compared with the state-of-the-art methods, especially in the case of only a few labeled samples and limited dictionary size. Hongteng Xu, Licheng Yu, Dixin Luo, Hongyuan Zha, Yi Xu 0001 |
AAAI | 2 |
| 2015 | Visual Madlibs: Fill in the Blank Description Generation and Question AnsweringabstractIn this paper, we introduce a new dataset consisting of 360,001 focused natural language descriptions for 10,738 images. This dataset, the Visual Madlibs dataset, is collected using automatically produced fill-in-the-blank templates designed to gather targeted descriptions about: people and objects, their appearances, activities, and interactions, as well as inferences about the general scene or its broader context. We provide several analyses of the Visual Madlibs dataset and demonstrate its applicability to two new description generation tasks: focused description generation, and multiple-choice question-answering for images. Experiments using joint-embedding and deep learning methods show promising results on these tasks. Licheng Yu, Eunbyung Park, Alexander C. Berg, Tamara L. Berg |
ICCV | 1 |
| 2015 | Analyzing Memory Access on CPU-GPGPU Shared LLC ArchitectureabstractThe data exchange between GPGPUs and CPUs are becoming more and more important nowadays. One trend in industry to alleviate the long latency is to integrate CPUs and GPGPUs on a single chip. In this paper, we analyze the reference interactions between CPU and GPGPU applications with a CPU-GPGPU co-simulator that integrates the gem5 and gpgpu-sim together. Since the memory controllers are shared among all cores, we observe severe memory contention between them. The CPU applications suffer a 1.26x slowdown and 64.79% blocked time in main memory when they run parallels with GPGPU applications. To alleviate the contention and provide more memory band-width, shared last level caches (LLCs) are commonly employed in such systems. We test a banked shared LLC structure that implanted into the co-simulator. We show that a simple shared LLC contributes mostly to the GPGPU (2.13x to running alone and 1.7x to running in parallel), rather than CPU. With the help of LLC, the memory requests issued to main memory is reduced to 30.74%, the blocked time is reduced to 49.64%, which provides more memory bandwidth. The latency-sensitive CPU applications are suffered as the LLC buffer occupation is very high when they run with GPGPU in parallel. Besides, as the number of LLC cache bank grows, we reveal that CPU achieves higher speedup than GPGPUs by increasing LLC parallelism. Finally, we also discuss the impact of GPGPU L2 cache. And we find that fewer GPGPU L2 cache banks will lower the performance as they limits the parallelism of GPGPU. The observations and inferences in this paper may serve as a reference guide to future CPU-GPGPU shared LLC design. Jianliang Ma, Licheng Yu, Tianzhou Chen, Minghui Wu 0001 |
ISPDC | 2 |
| 2015 | MCMG simulator: A unified simulation framework for CPU and graphic GPU
Jianliang Ma, Licheng Yu, John M. Ye, Tianzhou Chen |
J. Comput. Syst. Sci. | 2 |
| 2015 | Vector Sparse Representation of Color Image Using Quaternion Matrix AnalysisabstractTraditional sparse image models treat color image pixel as a scalar, which represents color channels separately or concatenate color channels as a monochrome image. In this paper, we propose a vector sparse representation model for color images using quaternion matrix analysis. As a new tool for color image representation, its potential applications in several image-processing tasks are presented, including color image reconstruction, denoising, inpainting, and super-resolution. The proposed model represents the color image as a quaternion matrix, where a quaternion-based dictionary learning algorithm is presented using the K-quaternion singular value decomposition (QSVD) (generalized K-means clustering for QSVD) method. It conducts the sparse basis selection in quaternion space, which uniformly transforms the channel images to an orthogonal color space. In this new color space, it is significant that the inherent color structures can be completely preserved during vector reconstruction. Moreover, the proposed sparse model is more efficient comparing with the current sparse models for image restoration tasks due to lower redundancy between the atoms of different color channels. The experimental results demonstrate that the proposed sparse image model avoids the hue bias issue successfully and shows its potential as a general and powerful tool in color image analysis and processing domain. Yi Xu 0001, Licheng Yu, Hongteng Xu, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2014 | Improving branch divergence performance on GPGPU with a new PDOM stack and multi-level warp scheduling
Licheng Yu, Xingsheng Tang, Minghui Wu 0001, Tianzhou Chen |
J. Syst. Archit. | 1 |
| 2013 | Quaternion-based sparse representation of color imageabstractIn this paper, we propose a quaternion-based sparse representation model for color images and its corresponding dictionary learning algorithm. Differing from traditional sparse image models, which represent RGB channels separately or process RGB channels as a concatenated real vector, the proposed model describes the color image as a quaternion vector matrix, where each color pixel is encoded as a quaternion unit and thus the inter-relationship among RGB channels is well preserved. Correspondingly, we propose a quaternion-based dictionary learning algorithm using a socalled K-QSVD method. It conducts the sparse basis selection in quaternion vector space, providing a kind of vectorial representation for the inherent color structures rather than a scalar representation via current sparse image models. The proposed sparse model is validated in the applications of color image denoising and inpainting. The experimental results demonstrate that our sparse image model avoids the hue bias phenomenon successfully and shows its potential as a powerful tool in color image analysis and processing domain. Licheng Yu, Yi Xu 0001, Hongteng Xu |
ICME | 1 |
| 2013 | Single image super-resolution via phase congruency analysisabstractSingle image super-resolution (SR) is a severely unconstrained task. While the self-example-based methods are able to reproduce sharp edges, they perform poorly for textures. For recovering the fine details, higher-level image segmentation and corresponding external texture database are employed in the example-based SR methods, but they involve too much human interaction. In this paper, we discuss the existing problems of example-based technique using scale space analysis. Accordingly, a robust pixel classification method is designed based on the phase congruency model in scale space, which can effectively divide images into edges, textures and flat regions. Then a super-resolution framework is proposed, which can adaptively emphasize the importance of high-frequency residuals in structural examples and scale invariant fractal property in textural regions. Experimental results show that our SR approach is able to present both sharp edges and vivid textures with few artifacts. Licheng Yu |
VCIP | 1 |
| 2012 | A CPU-GPGPU Scheduler Based on Data Transmission Bandwidth of WorkloadabstractWith the continuous development of GPUs, modern general-purpose computation on GPUs (GPGPUs) is providing growing parallelism to general programs besides graphics applications. However, for those programs that involve both CPU and GPU, the data transmission bandwidth between them may become bottleneck that prevents GPU from fully exploiting its parallel computing capacity. As to avoid the defect, we try to reduce the data transmission by keeping part of the computation tasks on the CPU side other than sending all the data over to the GPU and process there. In this way the computation is done on CPU and GPU in parallel, and therefore also reduces overall process time. In order to split the computation workload in a systematic approach, we try to divide the corresponding data into chunks of proper size. We experimented our data dividing and heterogeneous memory scheduling with 2 benchmarks. The matrix multiplication is more than 30% faster, and the k means2D is nearly 10% faster, than running solely in GPU. Licheng Yu, Minjiao Ye, Tianzhou Chen, Tongsen Hu |
PDCAT | 2 |
| 2012 | Packet Triggered Prediction Based Task Migration for Network-on-ChipabstractDeveloping IC technology makes Network-on-Chip (NoC) an attractive architecture for future systems. Task migration is important for the overall performance of NoCs since the changing system state makes static task mapping improper for NoCs. The predictability of behaviors of applications makes it possible to use prediction to guide task migration. The trigger to initiate task migration is also an important parameter. In this paper, we first defined and analyzed predictabilities of applications using experimental results. We also compared different triggers for migration and concluded that trigger based on packets sent by single node is the best choice. We modified Genetic Algorithm (GA) mapping for migration and proposed 2 algorithms Simple Exchange (SE) and Benefit Assess (BA). A Node Lock mechanism is also used to reduce the number of migrations. These algorithms and node lock are evaluated using real applications. According to the experimental results, SE reduced 78.7% of migrations with 9% less reduction of latency compared to GA; BA reduced 27.2% of latency, It reduced 72.0% of migrations with almost the same performance compared to GA. The node lock mechanism removed 37.3% and 46.0% of migrations in SE and BA with almost the same performance. Chao Wang 0058, Licheng Yu, Li Liu 0006, Tianzhou Chen |
PDP | 2 |
| 2011 | Leakage Aware Scheduling for Maximum Temperature MinimizationabstractAs power consumption continues to increase dramatically in real-time systems, the thermal management has become a prominent issue. Taking leakage current into account, this paper focuses on the maximum temperature minimization for the processor executing a set of real-time tasks with a common deadline. We prove that, for a specific interval, constant-speed schedule applying the lowest constant speed will be superior to any other schedule using higher constant speed in maximum temperature minimization. By dividing the interval into two subintervals, we develop a step-down scheduling algorithm, providing each subinterval a unique processor speed to further reduce the maximum temperature. Compared with the optimal constant-speed schedule, the proposed algorithm significantly reduces the maximum temperature by up to 12%. Jinming Yue, Tiefei Zhang, Licheng Yu, Tianzhou Chen |
PDCAT | 3 |