EDBT 2026 Demo / reviewers in the wild / expert
Jiuxiang Gu
dblp:173/4935
· DBLP profile ↗
60ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0002-3437-5084ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 10 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 28 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents ArchiveabstractThe opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and public policy. Analyzing how these interconnected systems simultaneously failed to protect public health requires innovative analytic approaches for exploring the vast amounts of data and documents disclosed in the UCSF-JHU Opioid Industry Documents Archive (OIDA). The complexity, multimodal nature, and specialized characteristics of these healthcare-related legal and corporate documents necessitate more advanced methods and models tailored to specific data types and detailed annotations, ensuring the precision and professionalism in the analysis. In this paper, we tackle this challenge by organizing the original dataset according to document attributes and constructing a benchmark with 400k training documents and 10k for testing. From each document, we extract rich multimodal information—including textual content, visual elements, and layout structures—to capture a comprehensive range of features. Using multiple AI models, we then generate a large-scale dataset comprising 360k training QA pairs and 10k testing QA pairs. Building on this foundation, we develop domain-specific multimodal Large Language Models (LLMs) and explore the impact of multimodal inputs on task performance. To further enhance response accuracy, we incorporate historical QA pairs as contextual grounding for answering current queries. Additionally, we incorporate page references within the answers and introduce an importance-based page classifier, further improving the precision and relevance of the information provided. Preliminary results indicate the improvements with our AI assistant in document information extraction and question-answering tasks. Xuan Shen, Brian Wingenroth, Zichao Wang 0001, Jason Kuen, Wanrong Zhu, Ruiyi Zhang 0002, Lichun Ma, Anqi Liu 0001, Tong Sun 0005, Kevin S. Hawkins, Kate Tasker, G. Caleb Alexander, Jiuxiang Gu |
AAAI | 15 |
| 2026 | VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-useabstractWhile vision-language models (VLMs) have demonstrated remarkable performance across various tasks combining textual and visual information, they continue to struggle with fine-grained visual perception tasks that require detailed pixel-level analysis. Effectively eliciting comprehensive reasoning from VLMs on such intricate visual elements remains an open challenge. In this paper, we present VipAct, an agent framework that enhances VLMs by integrating multi-agent collaboration and vision expert models, enabling more precise visual understanding and comprehensive reasoning. VipAct consists of an orchestrator agent, which manages task requirement analysis, planning, and coordination, along with specialized agents that handle specific tasks such as image captioning and vision expert models that provide high-precision perceptual information. This multi-agent approach allows VLMs to better perform fine-grained visual perception tasks by synergizing planning, reasoning, and tool use. We evaluate VipAct on benchmarks featuring a diverse set of visual perception tasks, with experimental results demonstrating significant performance improvements over state-of-the-art baselines across all tasks. Furthermore, comprehensive ablation studies reveal the critical role of multi-agent collaboration in eliciting more detailed System-2 reasoning and highlight the importance of image input for task planning. Additionally, our error analysis identifies patterns of VLMs' inherent limitations in visual perception, providing insights into potential future improvements. VipAct offers a flexible and extensible framework, paving the way for more advanced visual perception systems across various real-world applications. Zhehao Zhang 0001, Ryan Rossi, Tong Yu 0001, Franck Dernoncourt, Ruiyi Zhang 0002, Jiuxiang Gu, Sungchul Kim, Xiang Chen 0010, Zichao Wang 0001, Nedim Lipka |
AAAI | 6 |
| 2025 | LazyDiT: Lazy Learning for the Acceleration of Diffusion TransformersabstractDiffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications. The promising results come at the cost of slow inference, as each denoising step requires running the whole transformer model with a large amount of parameters. In this paper, we show that performing the full computation of the model at each diffusion step is unnecessary, as some computations can be skipped by lazily reusing the results of previous steps. Furthermore, we show that the lower bound of similarity between outputs at consecutive steps is notably high, and this similarity can be linearly approximated using the inputs. To verify our demonstrations, we propose the **LazyDiT**, a lazy learning framework that efficiently leverages cached results from earlier steps to skip redundant computations. Specifically, we incorporate lazy learning layers into the model, effectively trained to maximize laziness, enabling dynamic skipping of redundant computations. Experimental results show that LazyDiT outperforms the DDIM sampler across multiple diffusion transformer models at various resolutions. Furthermore, we implement our method on mobile devices, achieving better performance than DDIM with similar latency. Xuan Shen, Zhao Song 0002, Yufa Zhou 0001, Bo Chen 0029, Yanyu Li, Yifan Gong 0004, Kai Zhang 0045, Hao Tan 0002, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu 0002, Pu Zhao 0001, Yanzhi Wang 0001, Jiuxiang Gu |
AAAI | 15 |
| 2025 | Numerical Pruning for Efficient Autoregressive ModelsabstractTransformers have emerged as the leading architecture in deep learning, proving to be versatile and highly effective across diverse domains beyond language and image processing. However, their impressive performance often incurs high computational costs due to their substantial model size. This paper focuses on compressing decoder-only transformer-based autoregressive models through structural weight pruning to improve the model efficiency while preserving performance for both language and image generation tasks. Specifically, we propose a training-free pruning method that calculates a numerical score with Newton's method for the Attention and MLP modules, respectively. Besides, we further propose another compensation algorithm to recover the pruned model for better performance. To verify the effectiveness of our method, we provide both theoretical support and extensive experiments. Our experiments show that our method achieves state-of-the-art performance with reduced memory usage and faster generation speeds on GPUs. Xuan Shen, Zhao Song 0002, Yufa Zhou 0001, Bo Chen 0029, Jing Liu 0001, Ruiyi Zhang 0002, Ryan Rossi, Hao Tan 0005, Tong Yu 0001, Xiang Chen 0010, Yufan Zhou 0001, Tong Sun 0005, Pu Zhao 0001, Yanzhi Wang 0001, Jiuxiang Gu |
AAAI | 15 |
| 2025 | METAL: A Multi-Agent Framework for Chart Generation with Test-Time ScalingabstractChart generation aims to generate code to produce charts satisfying the desired visual properties, e.g., texts, layout, color, and type.It has great potential to empower the automatic professional report generation in financial analysis, research presentation, education, and healthcare.In this work, we build a vision-language model (VLM) based multi-agent framework for effective automatic chart generation.Generating high-quality charts requires both strong visual design skills and precise coding capabilities that embed the desired visual properties into code.Such a complex multi-modal reasoning process is difficult for direct prompting of VLMs.To resolve these challenges, we propose METAL (Multi-agEnT frAmework with vision Language models for chart generation), a multi-agent framework that decomposes the task of chart generation into the iterative collaboration among specialized agents.METAL achieves a 5.2% improvement in the F1 score over the current best result in the chart generation task.Additionally, METAL improves chart generation performance by 11.33% over Direct Prompting with LLAMA 3.2-11B.Furthermore, the METAL framework exhibits the phenomenon of test-time scaling: its performance increases monotonically as the logarithm of computational budget grows from 2 9 to 2 13 tokens. Yiwei Wang 0001, Jiuxiang Gu, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 3 |
| 2025 | From Selection to Generation: A Survey of LLM-based Active LearningabstractYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao, Nedim Lipka, Seunghyun Yoon, Ting-Hao Kenneth Huang, Zichao Wang, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee, Zhehao Zhang, Namyong Park, Thien Huu Nguyen, Jiebo Luo, Ryan A. Rossi, Julian McAuley. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Xia 0007, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li 0001, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen 0003, Franck Dernoncourt, Branislav Kveton, Tong Yu 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang 0160, Xiang Chen 0010, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao 0016, Nedim Lipka, Seunghyun Yoon 0002, Ting-Hao 'Kenneth' Huang, Zichao Wang 0001, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee 0001, Zhehao Zhang 0001, Namyong Park 0001, Thien Huu Nguyen, Jiebo Luo 0001, Ryan Rossi, Julian J. McAuley |
ACL (1) | 14 |
| 2025 | MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized DataabstractWe propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes—over 50 times larger than the prior real dataset DL3DV—dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth’s properties for enhancing model capability, training stability, and generalization, as well as application to other tasks. Hanwen Jiang, Zexiang Xu, Desai Xie, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang 0045, Sai Bi, Xin Sun 0014, Jiuxiang Gu, Qixing Huang, Georgios Pavlakos, Hao Tan 0002 |
CVPR | 11 |
| 2025 | QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the EdgeabstractMonocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous real-world applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the high computational and memory demands. Recent advancements in foundational depth estimation deliver impressive results but further amplify the difficulty of deployment on ASICs. To address this, we propose Quart-Depth which adopts post-training quantization to quantize MDE models with hardware accelerations for ASICs. Our approach involves quantizing both weights and activations to 4-bit precision, reducing the model size and computation cost. To mitigate the performance degradation, we introduce activation polishing and compensation algorithm applied before and after activation quantization, as well as a weight reconstruction method for minimizing errors in weight quantization. Furthermore, we design a flexible and programmable hardware accelerator by supporting kernel fusion and customized instruction programmability, enhancing throughput and efficiency. Experimental results demonstrate that our framework achieves competitive accuracy while enabling fast inference and higher energy efficiency on ASICs, bridging the gap between high-performance depth estimation and practical edge-device applicability. Code: https://github.com/shawnricecake/quart-depth Xuan Shen, Weize Ma, Jing Liu 0001, Changdi Yang, Quanyi Wang, Henghui Ding, Wei Niu 0002, Yanzhi Wang 0001, Pu Zhao 0001, Jiuxiang Gu |
CVPR | 12 |
| 2025 | CoMMIT: Coordinated Multimodal Instruction TuningabstractXintong Li, Junda Wu, Tong Yu, Rui Wang, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Julian McAuley, Jingbo Shang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xintong Li 0001, Junda Wu, Tong Yu 0001, Rui Wang 0088, Xiang Chen 0010, Jiuxiang Gu, Lina Yao 0001, Julian J. McAuley, Jingbo Shang |
EMNLP | 7 |
| 2025 | Refer to Any Segmentation Mask Group with Vision-Language PromptsabstractRecent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io. Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, Hyunjoon Jung, Liangyan Gui, Yu-Xiong Wang |
ICCV | 6 |
| 2025 | DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion Models
Zhuoling Li, Haoxuan Qu, Jason Kuen, Jiuxiang Gu, Qiuhong Ke, Jun Liu 0036, Hossein Rahmani 0001 |
ICCV | 4 |
| 2025 | Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
Shijie Zhou 0008, Ruiyi Zhang 0002, Huaisheng Zhu, Branislav Kveton, Yufan Zhou 0001, Jiuxiang Gu, Jian Chen 0043, Changyou Chen |
ICCV | 6 |
| 2025 | ImageFolder: Autoregressive Image Generation with Folded TokensabstractImage tokenizers are crucial for visual generative models, \eg, diffusion models (DMs) and autoregressive (AR) models, as they construct the latent representation for modeling. Increasing token length is a common approach to improve image reconstruction quality. However, tokenizers with longer token lengths are not guaranteed to achieve better generation quality. There exists a trade-off between reconstruction and generation quality regarding token length. In this paper, we investigate the impact of token length on both image reconstruction and generation and provide a flexible solution to the tradeoff. We propose \textbf{ImageFolder}, a semantic tokenizer that provides spatially aligned image tokens that can be folded during autoregressive modeling to improve both efficiency and quality. To enhance the representative capability without increasing token length, we leverage dual-branch product quantization to capture different contexts of images. Specifically, semantic regularization is introduced in one branch to encourage compacted semantic information while another branch is designed to capture pixel-level details. Extensive experiments demonstrate the superior quality of image generation and shorter token length with ImageFolder tokenizer. Xiang Li 0106, Hao Chen 0102, Jason Kuen, Jiuxiang Gu, Bhiksha Raj |
ICLR | 5 |
| 2025 | SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document UnderstandingabstractMultimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documents. Traditional methods using document parsers for retrieval-augmented generation suffer from performance and efficiency limitations, while directly presenting all pages to MLLMs leads to inefficiencies, especially with lengthy ones. In this work, we present a novel framework named **S**elf-**V**isual **R**etrieval-**A**ugmented **G**eneration (SV-RAG), which can broaden horizons of *any* MLLM to support long-document understanding. We demonstrate that **MLLMs themselves can be an effective multimodal retriever** to fetch relevant pages and then answer user questions based on these pages. SV-RAG is implemented with two specific MLLM adapters, one for evidence page retrieval and the other for question answering. Empirical results show state-of-the-art performance on public benchmarks, demonstrating the effectiveness of SV-RAG. Jian Chen 0043, Ruiyi Zhang 0002, Yufan Zhou 0001, Tong Yu 0001, Franck Dernoncourt, Jiuxiang Gu, Ryan Rossi, Changyou Chen, Tong Sun 0005 |
ICLR | 6 |
| 2025 | R-KV: Redundancy-aware KV Cache Compression for Reasoning ModelsabstractReasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 38% of the KV cache. This KV-cache reduction also leads to a 50% memory saving and a 2x speedup over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets. Zefan Cai, Hanshi Sun, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Anima Anandkumar, Abedelkadir Asi, Junjie Hu 0001 |
NeurIPS | 10 |
| 2025 | Differential Privacy Mechanisms in Neural Tangent Kernel RegressionabstractTraining data privacy is a fundamental problem in modern Artificial Intelligence (AI) applications, such as face recognition, recommendation systems, language generation, and many others, as it may contain sensitive user information related to legal issues. To fundamentally under-stand how privacy mechanisms work in AI applications, we study differential privacy (DP) in the Neural Tangent Kernel (NTK) regression setting, where DP is one of the most powerful tools for measuring privacy under statistical learning, and NTK is one of the most popular analysis frameworks for studying the learning mechanisms of deep neural networks. In our work, we can show provable guarantees for both differential privacy and test accuracy of our NTK regression. Furthermore, we conduct experiments on the basic image classification dataset CIFAR10 to demonstrate that NTK regression can preserve good accuracy under a modest privacy budget, supporting the validity of our analysis. To our knowledge, this is the first work to provide a DP guarantee for NTK regression. Jiuxiang Gu, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song 0002 |
WACV | 1 |
| 2025 | ARTIST: Improving the Generation of Text-Rich Images with Disentangled Diffusion Models and Large Language ModelsabstractDiffusion models have demonstrated exceptional capabilities in generating a broad spectrum of visual content, yet their proficiency in rendering text is still limited: they often generate inaccurate characters or words that fail to blend well with the underlying image. To address these shortcomings, we introduce a novel framework named ARTIST, which incorporates a dedicated textual diffusion model to focus on the learning of text structures specifically. Initially, we pretrain this textual model to capture the intricacies of text representation. Subsequently, we finetune a visual diffusion model, enabling it to assimilate textual structure information from the pretrained textual model. This disentangled architecture design and training strategy significantly enhance the text rendering ability of the diffusion models for text-rich image generation. Additionally, we leverage the capabilities of pretrained large language models to interpret user intentions better, contributing to improved generation quality. Empirical results on the MARIO-Eval benchmark underscore the effectiveness of the proposed method, showing an improvement of up to 15% in various metries. Yufan Zhou 0001, Jiuxiang Gu, Curtis Wigington, Tong Yu 0001, Yiran Chen 0001, Tong Sun 0005, Ruiyi Zhang 0002 |
WACV | 3 |
| 2024 | DocScript: Document-level Script Event PredictionabstractWe present a novel task of document-level script event prediction, which aims to predict the next event given a candidate list of narrative events in long-form documents. To enable this, we introduce DocSEP, a challenging dataset in two new domains - contractual documents and Wikipedia articles, where timeline events may be paragraphs apart and may require multi-hop temporal and causal reasoning. We benchmark existing baselines and present a novel architecture called DocScript to learn sequential ordering between events at the document scale. Our experimental results on the DocSEP dataset demonstrate that learning longer-range dependencies between events is a key challenge and show that contemporary LLMs such as ChatGPT and FlanT5 struggle to solve this task, indicating their lack of reasoning abilities for understanding causal relationships and temporal sequences within long texts. Puneet Mathur, Vlad I. Morariu, Aparna Garimella, Franck Dernoncourt, Jiuxiang Gu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha, Rajiv Jain |
LREC/COLING | 5 |
| 2024 | TRINS: Towards Multimodal Language Models that Can ReadabstractLarge multimodal language models have shown remarkable proficiency in understanding and editing images. However, a majority of these visually-tuned models struggle to comprehend the textual content embedded in images, primar-ily due to the limitation of training data. In this work, we introduce TRINS: a Text-Rich image11In this work, we use the phrase “text-rich images” to describe images with rich textual information, such as posters and book covers. INStruction dataset, with the objective of enhancing the reading ability of the multimodal large language model. TRINS is built upon LAION22Work done during Q3 2023. using hybrid data annotation strategies that include machine-assisted and human-assisted annotation process. It contains 39,153 text-rich images, captions, and 102,437 questions. Specifically, we show that the number of words per annotation in TRINS is significantly longer than that of related datasets, providing new challenges. Furthermore, we introduce a simple and effective architecture, called a Language-Vision Reading Assistant (LaRA), which is good at understanding textual content within images. LaRA outperforms existing state-of-the-art multimodal large language models on the TRINS dataset as well as other classical benchmarks. Lastly, we conducted a comprehensive evaluation with TRINS on various text-rich image understanding and generation tasks, demonstrating its effectiveness. Ruiyi Zhang 0002, Jian Chen 0043, Yufan Zhou 0001, Jiuxiang Gu, Changyou Chen, Tong Sun 0005 |
CVPR | 5 |
| 2024 | Customization Assistant for Text-to-image GenerationabstractCustomizing pre-trained text-to-image generation model has attracted massive research interest recently, due to its huge potential in real-world applications. Although existing methods are able to generate creative content for a novel concept contained in single user-input image, their capability are still far from perfection. Specifically, most existing methods require fine-tuning the generative model on testing images. Some existing methods do not require fine-tuning, while their performance are unsatisfactory. Furthermore, the interaction between users and models are still limited to directive and descriptive prompts such as instructions and captions. In this work, we build a customization assistant based on pre-trained large language model and diffusion model, which can not only perform customized generation in a tuning-free manner, but also enable more user-friendly interactions: users can chat with the assistant and input ei-ther ambiguous text or clear instruction. Specifically, we propose a new framework consists of a new model design and a novel training strategy. The resulting assistant can perform customized generation in 2–5 seconds without any test time fine-tuning. Extensive experiments are conducted, competitive results have been obtained across different domains, illustrating the effectiveness of the proposed method. Yufan Zhou 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Tong Sun 0005 |
CVPR | 3 |
| 2024 | SOHES: Self-supervised Open-world Hierarchical Entity SegmentationabstractOpen-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive generalization capabilities on unseen images and concepts. Despite its promise, existing entity segmentation methods like Segment Anything Model (SAM) rely heavily on costly expert annotators. This work presents Self-supervised Open-world Hierarchical Entity Segmentation (SOHES), a novel approach that eliminates the need for human annotations. SOHES operates in three phases: self-exploration, self-instruction, and self-correction. Given a pre-trained self-supervised representation, we produce abundant high-quality pseudo-labels through visual feature clustering. Then, we train a segmentation model on the pseudo-labels, and rectify the noises in pseudo-labels via a teacher-student mutual-learning procedure. Beyond segmenting entities, SOHES also captures their constituent parts, providing a hierarchical understanding of visual entities. Using raw images as the sole training data, our method achieves unprecedented performance in self-supervised open-world segmentation, marking a significant milestone towards high-quality open-world entity segmentation in the absence of human-annotated masks. Project page: https://SOHES.github.io. Shengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan 0002, Ruiyi Zhang 0002, Handong Zhao, Ani Nenkova, Liangyan Gui, Tong Sun 0005, Yu-Xiong Wang |
ICLR | 2 |
| 2024 | ADOPD: A Large-Scale Document Page Decomposition DatasetabstractResearch in document image understanding is hindered by limited high-quality document data. To address this, we introduce ADOPD, a comprehensive dataset for document page decomposition. ADOPD stands out with its data-driven approach for document taxonomy discovery during data collection, complemented by dense annotations. Our approach integrates large-scale pretrained models with a human-in-the-loop process to guarantee diversity and balance in the resulting data collection. Leveraging our data-driven document taxonomy, we collect and densely annotate document images, addressing four document image understanding tasks: Doc2Mask, Doc2Box, Doc2Tag, and Doc2Seq. Specifically, for each image, the annotations include human-labeled entity masks, text bounding boxes, as well as automatically generated tags and captions that have been manually cleaned. We conduct comprehensive experimental analyses to validate our data and assess the four tasks using various models. We envision ADOPD as a foundational dataset with the potential to drive future research in document understanding. Jiuxiang Gu, Xiangxi Shi, Jason Kuen, Ruiyi Zhang 0002, Anqi Liu 0001, Ani Nenkova, Tong Sun 0005 |
ICLR | 1 |
| 2024 | LRM: Large Reconstruction Model for Single Image to 3DabstractWe propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs, including real-world in-the-wild captures and images created by generative models. Video demos and interactable 3D meshes can be found on our LRM project webpage: https://yiconghong.me/LRM. Yicong Hong, Kai Zhang 0045, Jiuxiang Gu, Sai Bi, Yang Zhou 0009, Difan Liu, Feng Liu 0015, Kalyan Sunkavalli, Trung Bui, Hao Tan 0002 |
ICLR | 3 |
| 2024 | Category-Aware Active Domain AdaptationabstractActive domain adaptation has shown promising results in enhancing unsupervised domain adaptation (DA), by actively selecting and annotating a small amount of unlabeled samples from the target domain. Despite its effectiveness in boosting overall performance, the gain usually concentrates on the categories that are readily improvable, while challenging categories that demand the utmost attention are often overlooked by existing models. To alleviate this discrepancy, we propose a novel category-aware active DA method that aims to boost the adaptation for the individual category without adversely affecting others. Specifically, our approach identifies the unlabeled data that are most important for the recognition of the targeted category. Our method assesses the impact of each unlabeled sample on the recognition loss of the target data via the influence function, which allows us to directly evaluate the sample importance, without relying on indirect measurements used by existing methods. Comprehensive experiments and in-depth explorations demonstrate the efficacy of our method on category-aware active DA over three datasets. Wenxiao Xiao, Jiuxiang Gu, Hongfu Liu 0001 |
ICML | 2 |
| 2023 | DocEdit: Language-Guided Document EditingabstractProfessional document editing tools require a certain level of expertise to perform complex edit operations. To make editing tools accessible to increasingly novice users, we investigate intelligent document assistant systems that can make or suggest edits based on a user's natural language request. Such a system should be able to understand the user's ambiguous requests and contextualize them to the visual cues and textual content found in a document image to edit localized unstructured text and structured layouts. To this end, we propose a new task of language-guided localized document editing, where the user provides a document and an open vocabulary editing request, and the intelligent system produces a command that can be used to automate edits in real-world document editing software. In support of this task, we curate the DocEdit dataset, a collection of approximately 28K instances of user edit requests over PDF and design templates along with their corresponding ground truth software executable commands. To our knowledge, this is the first dataset that provides a diverse mix of edit operations with direct and indirect references to the embedded text and visual objects such as paragraphs, lists, tables, etc. We also propose DocEditor, a Transformer-based localization-aware multimodal (textual, spatial, and visual) model that performs the new task. The model attends to both document objects and related text contents which may be referred to in a user edit request, generating a multimodal embedding that is used to predict an edit command and associated bounding box localizing it. Our proposed model empirically outperforms other baseline deep learning approaches by 15-18%, providing a strong starting point for future work. Puneet Mathur, Rajiv Jain, Jiuxiang Gu, Franck Dernoncourt, Dinesh Manocha, Vlad I. Morariu |
AAAI | 3 |
| 2023 | Learning the Visualness of Text Using Large Vision-Language ModelsabstractVisual text evokes an image in a person's mind, while non-visual text fails to do so.A method to automatically detect visualness in text will enable text-to-image retrieval and generation models to augment text with relevant images.This is particularly challenging with long-form text as text-to-image generation and retrieval models are often triggered for text that is designed to be explicitly visual in nature, whereas long-form text could contain many non-visual sentences.To this end, we curate a dataset of 3,620 English sentences and their visualness scores provided by multiple human annotators.We also propose a fine-tuning strategy that adapts large vision-language models like CLIP by modifying the model's contrastive learning objective to map text identified as nonvisual to a common NULL image while matching visual text to their corresponding images in the document.We evaluate the proposed approach on its ability to (i) classify visual and non-visual text accurately, and (ii) attend over words that are identified as visual in psycholinguistic studies.Empirical evaluation indicates that our approach performs better than several heuristics and baseline models for the proposed task.Furthermore, to highlight the importance of modeling the visualness of text, we conduct qualitative analyses of text-to-image generation systems like DALL-E. Gaurav Verma 0005, Ryan Rossi, Chris Tensmeyer, Jiuxiang Gu, Ani Nenkova |
EMNLP | 4 |
| 2023 | High Quality Entity SegmentationabstractDense image segmentation tasks (e.g., semantic, panoptic) are useful for image editing, but existing methods can hardly generalize well in an in-the-wild setting where there are unrestricted image domains, classes, and image resolution & quality variations. Motivated by these observations, we construct a new entity segmentation dataset, with a strong focus on high-quality dense segmentation in the wild. The dataset contains images spanning diverse image domains and entities, along with plentiful high-resolution images and high-quality mask annotations for training and testing. Given the high-quality and -resolution nature of the dataset, we propose CropFormer which is designed to tackle the intractability of instance-level segmentation on high-resolution images. It improves mask prediction by fusing high-res image crops that provides more fine-grained image details and the full image. CropFormer is the first query-based Transformer architecture that can effectively fuse mask predictions from multiple image views, by learning queries that effectively associate the same entities across the full image and its crop. With CropFormer, we achieve a significant AP gain of 1.9 on the challenging entity segmentation task. Furthermore, CropFormer consistently improves the accuracy of traditional segmentation tasks and datasets. The dataset and code are released at http://luqi.info/entityv2.github.io/. Lu Qi 0001, Jason Kuen, Tiancheng Shen, Jiuxiang Gu, Wenbo Li 0001, Weidong Guo, Jiaya Jia, Zhe Lin 0001, Ming-Hsuan Yang 0001 |
ICCV | 4 |
| 2023 | AIMS: All-Inclusive Multi-Level Segmentation for AnythingabstractDespite the progress of image segmentation for accurate visual entity segmentation, completing the diverse requirements of image editing applications for different-level region-of-interest selections remains unsolved. In this paper, we propose a new task, All-Inclusive Multi-Level Segmentation (AIMS), which segments visual regions into three levels: part, entity, and relation (two entities with some semantic relationships). We also build a unified AIMS model through multi-dataset multi-task training to address the two major challenges of annotation inconsistency and task correlation. Specifically, we propose task complementarity, association, and prompt mask encoder for three-level predictions. Extensive experiments demonstrate the effectiveness and generalization capacity of our method compared to other state-of-the-art methods on a single dataset or the concurrent work on segment anything. We will make our code and training model publicly available. Lu Qi 0001, Jason Kuen, Weidong Guo, Jiuxiang Gu, Zhe Lin 0001, Bo Du 0001, Ming-Hsuan Yang 0001 |
NeurIPS | 4 |
| 2023 | LayerDoc: Layer-wise Extraction of Spatial Hierarchical Structure in Visually-Rich DocumentsabstractDigital documents often contain images and scanned text. Parsing such visually-rich documents is a core task for work-flow automation, but it remains challenging since most documents do not encode explicit layout information, e.g., how characters and words are grouped into boxes and ordered into larger semantic entities. Current state-of-the-art layout extraction methods are challenged by such documents as they rely on word sequences to have correct reading order and do not exploit their hierarchical structure. We propose LayerDoc, an approach that uses visual features, textual semantics, and spatial coordinates along with constraint inference to extract the hierarchical layout structure of documents in a bottom-up layer-wise fashion. LayerDoc recursively groups smaller regions into larger semantic elements in 2D to infer complex nested hierarchies. Experiments show that our approach outperforms competitive baselines by 10-15% on three diverse datasets of forms and mobile app screen layouts for the tasks of spatial region classification, higher-order group identification, layout hierarchy extraction, reading order detection, and word grouping. Puneet Mathur, Rajiv Jain, Ashutosh Mehra 0002, Jiuxiang Gu, Franck Dernoncourt, Anandhavelu Natarajan, Quan Hung Tran, Verena Kaynig, Ani Nenkova, Dinesh Manocha, Vlad I. Morariu |
WACV | 4 |
| 2023 | Open World Entity SegmentationabstractWe introduce a new image segmentation task, called Entity Segmentation (ES), which aims to segment all visual entities (objects and stuffs) in an image without predicting their semantic labels. By removing the need of class label prediction, the models trained for such task can focus more on improving segmentation quality. It has many practical applications such as image manipulation and editing where the quality of segmentation masks is crucial but class labels are less important. We conduct the first-ever study to investigate the feasibility of convolutional center-based representation to segment things and stuffs in a unified manner, and show that such representation fits exceptionally well in the context of ES. More specifically, we propose a CondInst-like fully-convolutional architecture with two novel modules specifically designed to exploit the class-agnostic and non-overlapping requirements of ES. Experiments show that the models designed and trained for ES significantly outperforms popular class-specific panoptic segmentation models in terms of segmentation quality. Moreover, an ES model can be easily trained on a combination of multiple datasets without the need to resolve label conflicts in dataset merging, and the model trained for ES on one or more datasets can generalize very well to other test datasets of unseen domains. The code has been released at https://github.com/dvlab-research/Entity. Lu Qi 0001, Jason Kuen, Yi Wang 0074, Jiuxiang Gu, Hengshuang Zhao, Philip Torr 0001, Zhe Lin 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | UNISON: Unpaired Cross-Lingual Image CaptioningabstractImage captioning has emerged as an interesting research field in recent years due to its broad application scenarios. The traditional paradigm of image captioning relies on paired image-caption datasets to train the model in a supervised manner. However, creating such paired datasets for every target language is prohibitively expensive, which hinders the extensibility of captioning technology and deprives a large part of the world population of its benefit. In this work, we present a novel unpaired cross-lingual method to generate image captions without relying on any caption corpus in the source or the target language. Specifically, our method consists of two phases: (1) a cross-lingual auto-encoding process, which utilizing a sentence parallel (bitext) corpus to learn the mapping from the source to the target language in the scene graph encoding space and decode sentences in the target language, and (2) a cross-modal unsupervised feature mapping, which seeks to map the encoded scene graph features from image modality to language modality. We verify the effectiveness of our proposed method on the Chinese image caption generation task. The comparisons against several existing methods demonstrate the effectiveness of our approach. Jiahui Gao 0002, Yi Zhou 0042, Philip L. H. Yu, Shafiq R. Joty, Jiuxiang Gu |
AAAI | 5 |
| 2022 | TiGAN: Text-Based Interactive Image Generation and ManipulationabstractUsing natural-language feedback to guide image generation and manipulation can greatly lower the required efforts and skills. This topic has received increased attention in recent years through refinement of Generative Adversarial Networks (GANs); however, most existing works are limited to single-round interaction, which is not reflective of real world interactive image editing workflows. Furthermore, previous works dealing with multi-round scenarios are limited to predefined feedback sequences, which is also impractical. In this paper, we propose a novel framework for Text-based Interactive image generation and manipulation (TiGAN) that responds to users' natural-language feedback. TiGAN utilizes the powerful pre-trained CLIP model to understand users' natural-language feedback and exploits contrastive learning for a better text-to-image mapping. To maintain the image consistency during interactions, TiGAN generates intermediate feature vectors aligned with the feedback and selectively feeds these vectors to our proposed generative model. Empirical results on several datasets show that TiGAN improves both interaction efficiency and image quality while better avoids undesirable image manipulation during interactions. Yufan Zhou 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Chris Tensmeyer, Tong Yu 0001, Changyou Chen, Jinhui Xu 0001, Tong Sun 0005 |
AAAI | 3 |
| 2022 | User-Entity Differential Privacy in Learning Natural Language ModelsabstractIn this paper, we introduce a novel concept of user-entity differential privacy (UeDP) to provide formal privacy protection simultaneously to both sensitive entities in textual data and data owners in learning natural language models (NLMs). To preserve UeDP, we developed a novel algorithm, called UeDP-Alg, optimizing the trade-off between privacy loss and model utility with a tight sensitivity bound derived from seamlessly combining user and sensitive entity sampling processes. An extensive theoretical analysis and evaluation show that our UeDP-Alg outperforms baseline approaches in model utility under the same privacy budget consumption on several NLM tasks, using benchmark datasets. Phung Lai, NhatHai Phan, Tong Sun 0005, Rajiv Jain, Franck Dernoncourt, Jiuxiang Gu, Nikolaos Barmpalios |
IEEE Big Data | 6 |
| 2022 | Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-LabelingabstractOpen-vocabulary instance segmentation aims at segmenting novel classes without mask annotations. It is an important step toward reducing laborious human supervision. Most existing works first pretrain a model on captioned images covering many novel classes and then finetune it on limited base classes with mask annotations. However, the high-level textual information learned from caption pretraining alone cannot effectively encode the details required for pixelwise segmentation. To address this, we propose a cross-modal pseudo-labeling framework, which generates training pseudo masks by aligning word semantics in captions with visual features of object masks in images. Thus, our framework is capable of labeling novel classes in captions via their word semantics to self-train a student model. To account for noises in pseudo masks, we design a robust student model that selectively distills mask knowledge by estimating the mask noise levels, hence mitigating the adverse impact of noisy pseudo masks. By extensive experiments, we show the effectiveness of our framework, where we significantly improve mAP score by 4.5% on MS-COCO and 5.1 % on the large-scale Open Images & Conceptual Captions datasets compared to the state-of-the-art.11Code is available at https://github.com/hbdat/cvpr22_cross_modal_pseudo_labeling. Dat Huynh, Jason Kuen, Zhe Lin 0001, Jiuxiang Gu, Ehsan Elhamifar |
CVPR | 4 |
| 2022 | EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal RetrievalabstractCross language-image modality retrieval in E-commerce is a fundamental problem for product search, recommendation, and marketing services. Extensive efforts have been made to conquer the cross-modal retrieval problem in the general domain. When it comes to E-commerce, a com-mon practice is to adopt the pretrained model and finetune on E-commerce data. Despite its simplicity, the performance is sub-optimal due to overlooking the uniqueness of E-commerce multimodal data. A few recent efforts [10], [72] have shown significant improvements over generic methods with customized designs for handling product images. Unfortunately, to the best of our knowledge, no existing method has addressed the unique challenges in the e-commerce language. This work studies the outstanding one, where it has a large collection of special meaning entities, e.g., “Di s s e l (brand)”, “Top (category)”, “relaxed (fit)” in the fashion clothing business. By formulating such out-of-distribution finetuning process in the Causal Inference paradigm, we view the erroneous semantics of these special entities as confounders to cause the retrieval failure. To rectify these semantics for aligning with e-commerce do-main knowledge, we propose an intervention-based entity-aware contrastive learning framework with two modules, i.e., the Confounding Entity Selection Module and Entity-Aware Learning Module. Our method achieves competitive performance on the E-commerce benchmark Fashion-Gen. Particularly, in top-1 accuracy (R@l), we observe 10.3% and 10.5% relative improvements over the closest baseline in image-to-text and text-to-image retrievals, respectively. Handong Zhao, Zhe Lin 0001, Ajinkya Kale, Zhangyang Wang, Tong Yu 0001, Jiuxiang Gu, Sunav Choudhary, Xiaohui Xie |
CVPR | 7 |
| 2022 | Towards Language-Free Training for Text-to-Image GenerationabstractOne of the major challenges in training text-to-image generation models is the need of a large number of highquality image-text pairs. While image samples are often easily accessible, the associated text descriptions typically require careful human captioning, which is particularly time- and cost-consuming. In this paper, we propose the first work to train text-to-image generation models without any text data. Our method leverages the well-aligned multi-modal semantic space of the powerful pre-trained CLIP model: the requirement of text-conditioning is seamlessly alleviated via generating text features from image features. Extensive experiments are conducted to illustrate the effectiveness of the proposed method. We obtain state-of-the-art results in the standard text-to-image generation tasks. Importantly, the proposed language-free model outperforms most existing models trained with full image-text pairs. Furthermore, our method can be applied in fine-tuning pretrained models, which saves both training time and cost in training text-to-image generation models. Our pre-trained model obtains competitive results in zero-shot text-to-image generation on the MS-COCO dataset, yet with around only 1% of the model size and training data size relative to the recently proposed large DALL-E model. Yufan Zhou 0001, Ruiyi Zhang 0002, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu 0001, Jiuxiang Gu, Jinhui Xu 0001, Tong Sun 0005 |
CVPR | 7 |
| 2022 | CA-SSL: Class-Agnostic Semi-Supervised Learning for Detection and Segmentation
Lu Qi 0001, Jason Kuen, Zhe Lin 0001, Jiuxiang Gu, Fengyun Rao, Weidong Guo, Ming-Hsuan Yang 0001, Jiaya Jia |
ECCV (31) | 4 |
| 2022 | Improving the Reliability for Confidence Estimation
Haoxuan Qu, Lin Geng Foo, Jason Kuen, Jiuxiang Gu, Jun Liu 0036 |
ECCV (27) | 5 |
| 2022 | Meta Spatio-Temporal Debiasing for Video Scene Graph Generation
Haoxuan Qu, Jason Kuen, Jiuxiang Gu, Jun Liu 0036 |
ECCV (27) | 4 |
| 2022 | MGDoc: Pre-training with Multi-granular Hierarchy for Document Image UnderstandingabstractZilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun, Jingbo Shang, Vlad Morariu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Zilong Wang 0002, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005, Jingbo Shang, Vlad I. Morariu |
EMNLP | 2 |
| 2022 | DocLayoutTTS: Dataset and Baselines for Layout-informed Document-level Neural Speech Synthesis
Puneet Mathur, Franck Dernoncourt, Quan Hung Tran, Jiuxiang Gu, Ani Nenkova, Vlad I. Morariu, Rajiv Jain, Dinesh Manocha |
INTERSPEECH | 4 |
| 2022 | DocTime: A Document-level Temporal Dependency Graph ParserabstractPuneet Mathur, Vlad Morariu, Verena Kaynig-Fittkau, Jiuxiang Gu, Franck Dernoncourt, Quan Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Puneet Mathur, Vlad I. Morariu, Verena Kaynig, Jiuxiang Gu, Franck Dernoncourt, Quan Hung Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain |
NAACL-HLT | 4 |
| 2022 | Delving into Out-of-Distribution Detection with Vision-Language RepresentationsabstractRecognizing out-of-distribution (OOD) samples is critical for machine learning systems deployed in the open world. The vast majority of OOD detection methods are driven by a single modality (e.g., either vision or language), leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of OOD detection from a single-modal to a multi-modal regime. Particularly, we propose Maximum Concept Matching (MCM), a simple yet effective zero-shot OOD detection method based on aligning visual features with textual concepts. We contribute in-depth analysis and theoretical insights to understand the effectiveness of MCM. Extensive experiments demonstrate that MCM achieves superior performance on a wide variety of real-world tasks. MCM with vision-language features outperforms a common baseline with pure visual features on a hard OOD task with semantically similar classes by 13.1% (AUROC) Code is available at https://github.com/deeplearning-wisc/MCM. Yifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun, Yixuan Li 0001 |
NeurIPS | 3 |
| 2022 | FedKC: Federated Knowledge Composition for Multilingual Natural Language UnderstandingabstractMultilingual natural language understanding, which aims to comprehend multilingual documents, is an important task. Existing efforts have been focusing on the analysis of centrally stored text data, but in real practice, multilingual data is usually distributed. Federated learning is a promising paradigm to solve this problem, which trains local models with decentralized data on local clients and aggregates local models on the central server to achieve a good global model. However, existing federated learning methods assume that data are independent and identically distributed (IID), and cannot handle multilingual data, that are usually non-IID with severely skewed distributions: First, multilingual data is stored on local client devices such that there are only monolingual or bilingual data stored on each client. This makes it difficult for local models to know the information of documents in other languages. Second, the distribution over different languages could be skewed. High resource language data is much more abundant than low resource language data. The model trained on such skewed data may focus more on high resource languages but fail to consider the key information of low resource languages. To solve the aforementioned challenges of multilingual federated NLU, we propose a plug-and-play knowledge composition (KC) module, called FedKC, which exchanges knowledge among clients without sharing raw data. Specifically, we propose an effective way to calculate a consistency loss defined based on the shared knowledge across clients, which enables models trained on different clients achieve similar predictions on similar data. Leveraging this consistency loss, joint training is thus conducted on distributed data respecting the privacy constraints. We also analyze the potential risk of FedKC and provide theoretical bound to show that it is difficult to recover data from the corrupted data. We conduct extensive experiments on three public multilingual datasets for three typical NLU tasks, including paraphrase identification, question answering matching, and news classification. The experiment results show that the proposed FedKC can outperform state-of-the-art baselines on the three datasets significantly. Haoyu Wang 0004, Handong Zhao, Yaqing Wang 0001, Tong Yu 0001, Jiuxiang Gu, Jing Gao 0004 |
WWW | 5 |
| 2021 | SelfDoc: Self-Supervised Document Representation LearningabstractWe propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and visual information of every semantically meaningful component in a document, and it models the contextualization between each block of content. Unlike existing document pre-training models, our model is coarse-grained instead of treating individual words as input, therefore avoiding an overly fine-grained with excessive contextualization. Beyond that, we introduce cross-modal learning in the model pre-training phase to fully leverage multimodal information from unlabeled documents. For downstream usage, we propose a novel modality-adaptive attention mechanism for multimodal feature fusion by adaptively emphasizing language and vision signals. Our framework benefits from self-supervised pre-training on documents without requiring annotations by a feature masking training strategy. It achieves superior performance on multiple downstream tasks with significantly fewer document images used in the pre-training stage compared to previous works. Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, Hongfu Liu 0001 |
CVPR | 2 |
| 2021 | Multi-Scale Aligned Distillation for Low-Resolution DetectionabstractIn instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option traditionally hurts the detection performance much. This paper focuses on boosting performance of low-resolution models by distilling knowledge from a high- or multi-resolution model. We first identify the challenge of applying knowledge distillation (KD) to teacher and student networks that act on different input resolutions. To tackle it, we explore the idea of spatially aligning feature maps between models of varying input resolutions by shifting feature pyramid position and introduce aligned multi-scale training to train a multi-scale teacher that can distill its knowledge to a low-resolution student. Further, we propose crossing feature-level fusion to dynamically fuse teacher’s multi-resolution features to guide the student better. On several instance-level detection tasks and datasets, the low-resolution models trained via our approach perform competitively with high-resolution models trained via conventional multi-scale training, while outperforming the latter’s low-resolution models by 2.1% to 3.6% in terms of mAP. Our code is made publicly available at https://github.com/Jia-Research-Lab/MSAD. Lu Qi 0001, Jason Kuen, Jiuxiang Gu, Zhe Lin 0001, Yi Wang 0074, Yukang Chen, Jiaya Jia |
CVPR | 3 |
| 2021 | Exploiting Semantic Embedding and Visual Feature for Facial Action Unit DetectionabstractRecent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection performance. As of now, no semantic information of AUs has yet been explored for such a task. As a matter of fact, AU semantic descriptions provide much more information than the binary AU labels alone, thus we propose to exploit the Semantic Embedding and Visual feature (SEV-Net) for AU detection. More specifically, AU semantic embeddings are obtained through both Intra-AU and Inter-AU attention modules, where the Intra-AU attention module captures the relation among words within each sentence that describes individual AU, and the Inter-AU attention module focuses on the relation among those sentences. The learned AU semantic embeddings are then used as guidance for the generation of attention maps through a cross-modality attention network. The generated cross-modality attention maps are further used as weights for the aggregated feature. Our proposed method is unique in that the semantic features are exploited as the first of this kind. The approach has been evaluated on three public AU-coded facial expression databases, and has achieved a superior performance than the state-of-the-art peer methods. Huiyuan Yang, Lijun Yin 0001, Jiuxiang Gu |
CVPR | 4 |
| 2021 | Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU modelsabstractMengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, Xia Hu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun 0005, Xia Ben Hu |
NAACL-HLT | 6 |
| 2021 | UniDoc: Unified Pretraining Framework for Document UnderstandingabstractDocument intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up promising directions towards reducing annotation efforts by training models with self-supervised objectives. However, most of the existing document pretraining methods are still language-dominated. We present UDoc, a new unified pretraining framework for document understanding. UDoc is designed to support most document understanding tasks, extending the Transformer to take multimodal embeddings as input. Each input element is composed of words and visual features from a semantic region of the input document image. An important feature of UDoc is that it learns a generic representation by making use of three self-supervised losses, encouraging the representation to model sentences, learn similarities, and align modalities. Extensive empirical analysis demonstrates that the pretraining procedure learns better joint representations and leads to improvements in downstream tasks. Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005 |
NeurIPS | 1 |
| 2020 | Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning
Xiangxi Shi, Xu Yang 0021, Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001 |
ECCV (14) | 3 |
| 2020 | Self-Supervised Relationship ProbingabstractStructured representations of images that model visual relationships are beneficial for many vision and vision-language applications. However, current human-annotated visual relationship datasets suffer from the long-tailed predicate distribution problem which limits the potential of visual relationship models. In this work, we introduce a self-supervised method that implicitly learns the visual relationships without relying on any ground-truth visual relationship annotations. Our method relies on 1) intra- and inter-modality encodings to respectively model relationships within each modality separately and jointly, and 2) relationship probing, which seeks to discover the graph structure within each modality. By leveraging masked language modeling, contrastive learning, and dependency tree distances for self-supervision, our method learns better object features as well as implicit visual relationships. We verify the effectiveness of our proposed method on various vision-language tasks that benefit from improved visual relationship understanding. Jiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai 0001, Vlad I. Morariu, Handong Zhao, Tong Sun 0005 |
NeurIPS | 1 |
| 2020 | Video captioning with boundary-aware hierarchical language decoding and joint video prediction
Xiangxi Shi, Jianfei Cai 0001, Jiuxiang Gu, Shafiq R. Joty |
Neurocomputing | 3 |
| 2019 | Scene Graph Generation With External Knowledge and Image ReconstructionabstractScene graph generation has received growing attention with the advancements in image understanding tasks such as object detection, attributes and relationship prediction, etc. However, existing datasets are biased in terms of object and relationship labels, or often come with noisy and missing annotations, which makes the development of a reliable scene graph prediction model very challenging. In this paper, we propose a novel scene graph generation algorithm with external knowledge and image reconstruction loss to overcome these dataset issues. In particular, we extract commonsense knowledge from the external knowledge base to refine object and phrase features for improving generalizability in scene graph generation. To address the bias of noisy object annotations, we introduce an auxiliary image reconstruction path to regularize the scene graph generation network. Extensive experiments show that our framework can generate better scene graphs, achieving the state-of-the-art performance on two benchmark datasets: Visual Relationship Detection and Visual Genome datasets. Jiuxiang Gu, Handong Zhao, Zhe Lin 0001, Sheng Li 0001, Jianfei Cai 0001, Mingyang Ling 0001 |
CVPR | 1 |
| 2019 | Unpaired Image Captioning via Scene Graph AlignmentsabstractMost of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming. In this paper, we present a scene graph-based approach for unpaired image captioning. Our framework comprises an image scene graph generator, a sentence scene graph generator, a scene graph encoder, and a sentence decoder. Specifically, we first train the scene graph encoder and the sentence decoder on the text modality. To align the scene graphs between images and sentences, we propose an unsupervised feature alignment method that maps the scene graph features from the image to the sentence modality. Experimental results show that our proposed model can generate quite promising results without using any image-caption training pairs, outperforming existing methods by a wide margin. Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001, Handong Zhao, Xu Yang 0021, Gang Wang 0012 |
ICCV | 1 |
| 2019 | Watch It Twice: Video Captioning with a Refocused Video EncoderabstractWith the rapid growth of video data and the increasing demands of various crossmodal applications such as intelligent video search and assistance towards visually-impaired people, video captioning task has received a lot of attention recently in computer vision and natural language processing fields. The state-of-the-art video captioning methods focus more on encoding the temporal information, while lacking effective ways to remove irrelevant temporal information and also neglecting the spatial details. In particular, the current unidirectional video encoder can be negatively affected by irrelevant temporal information, especially the irrelevant information at the beginning and at the end of a video. In addition, disregarding detailed spatial features may lead to incorrect word choices in decoding. In this paper, we propose a novel recurrent video encoding method and a novel visual spatial feature for the video captioning task. The recurrent encoding module encodes the video twice with a predicted key frame to avoid irrelevant temporal information often occurring at the beginning and at the end of a video. The novel spatial features represent spatial information from different regions of a video and provide the decoder with more detailed information. Experiments on two benchmark datasets show superior performance of the proposed method. Xiangxi Shi, Jianfei Cai 0001, Shafiq R. Joty, Jiuxiang Gu |
ACM Multimedia | 4 |
| 2018 | Stack-Captioning: Coarse-to-Fine Learning for Image CaptioningabstractThe existing image captioning approaches typically train a one-stage sentence decoder, which is difficult to generate rich fine-grained descriptions. On the other hand, multi-stage image caption model is hard to train due to the vanishing gradient problem. In this paper, we propose a coarse-to-fine multi-stage prediction framework for image captioning, composed of multiple decoders each of which operates on the output of the previous stage, producing increasingly refined image descriptions. Our proposed learning approach addresses the difficulty of vanishing gradients during training by providing a learning objective function that enforces intermediate supervisions. Particularly, we optimize our model with a reinforcement learning approach which utilizes the output of each intermediate decoder's test-time inference algorithm as well as the output of its preceding decoder to normalize the rewards, which simultaneously solves the well-known exposure bias problem and the loss-evaluation mismatch problem. We extensively evaluate the proposed approach on MSCOCO and show that our approach can achieve the state-of-the-art performance. Jiuxiang Gu, Jianfei Cai 0001, Gang Wang 0012, Tsuhan Chen |
AAAI | 1 |
| 2018 | Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval With Generative ModelsabstractTextual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that embed image-text pairs as single feature vectors in a common representational space, we propose to incorporate generative processes into the cross-modal feature embedding, through which we are able to learn not only the global abstract features but also the local grounded features. Extensive experiments show that our framework can well match images and sentences with complex content, and achieve the state-of-the-art cross-modal retrieval results on MSCOCO dataset. Jiuxiang Gu, Jianfei Cai 0001, Shafiq R. Joty, Li Niu 0002, Gang Wang 0012 |
CVPR | 1 |
| 2018 | Unpaired Image Captioning by Language Pivoting
Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001, Gang Wang 0012 |
ECCV (1) | 1 |
| 2018 | Recent advances in convolutional neural networks
Jiuxiang Gu, Zhenhua Wang 0002, Jason Kuen, Lianyang Ma, Amir Shahroudy, Bing Shuai, Ting Liu 0009, Gang Wang 0012, Jianfei Cai 0001, Tsuhan Chen |
Pattern Recognit. | 1 |
| 2017 | An Empirical Study of Language CNN for Image CaptioningabstractLanguage models based on recurrent neural networks have dominated recent image caption generation tasks. In this paper, we introduce a language CNN model which is suitable for statistical language modeling tasks and shows competitive performance in image captioning. In contrast to previous models which predict next word based on one previous word and hidden state, our language CNN is fed with all the previous words and can model the long-range dependencies in history words, which are critical for image captioning. The effectiveness of our approach is validated on two datasets: Flickr30K and MS COCO. Our extensive experimental results show that our method outperforms the vanilla recurrent neural network based language models and is competitive with the state-of-the-art methods. Jiuxiang Gu, Gang Wang 0012, Jianfei Cai 0001, Tsuhan Chen |
ICCV | 1 |