EDBT 2026 Demo / reviewers in the wild / expert
Liwei Wang 0009
dblp:47/1798-9
· DBLP profile ↗
47ranked-venue papers
6as first author
29since 2021 · last 2025
0000-0003-3264-1294ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning to Reason from Feedback at Test-TimeabstractSolving complex tasks in a single attempt is challenging for large language models (LLMs). Iterative interaction with the environment and feedback is often required to achieve success, making effective feedback utilization a critical topic. Existing approaches either struggle with length generalization or rely on naive retries without leveraging prior information. In this paper, we introduce FTTT, a novel paradigm that formulates feedback utilization as an optimization problem at test time. Additionally, we propose a learnable test-time optimizer, OpTune, to effectively exploit feedback. Experiments on two LLMs across four reasoning datasets demonstrate that FTTT and OpTune achieve superior scalability and performance. Yanyang Li, Michael R. Lyu, Liwei Wang 0009 |
ACL (1) | 3 |
| 2025 | Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene UnderstandingabstractThe rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to enhance MLLMs, such as incorporating point cloud features, have been made, yet a considerable gap remains between the models’ learned representations and the inherent complexity of 3D scenes. This discrepancy largely stems from the training of MLLMs on predominantly 2D data, which restricts their effectiveness in comprehending 3D spaces. To address this issue, in this paper, we propose a novel generalist model, i.e., Video-3D LLM, for 3D scene understanding. By treating 3D scenes as dynamic videos and incorporating 3D position encoding into these representations, our Video-3D LLM aligns video representations with real-world spatial contexts more accurately. In addition, we have implemented a maximum coverage sampling technique to optimize the trade-off between computational cost and performance. Extensive experiments demonstrate that our model achieves state-of-the-art performance on several 3D scene understanding benchmarks, including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D. Our code is available at https://github.com/LaVi-Lab/Video-3D-LLM. Duo Zheng, Shijia Huang, Liwei Wang 0009 |
CVPR | 3 |
| 2025 | Fine-Grained Spatiotemporal Grounding on Egocentric VideosabstractSpatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite its growing importance in applications such as augmented reality and robotics. In this work, we conduct a systematic analysis of the discrepancies between egocentric and exocentric videos, revealing key challenges such as shorter object durations, sparser trajectories, smaller object sizes, and larger positional shifts. To address these challenges, we introduce EgoMask, the first pixel-level benchmark for fine-grained spatiotemporal grounding in egocentric videos. It is constructed by our proposed automatic annotation pipeline, which annotates referring expressions and object masks across short-, medium-, and long-term videos. Additionally, we create EgoMask-Train, a large-scale training dataset to facilitate model development. Experiments demonstrate that the state-of-the-art spatiotemporal grounding models perform poorly on our benchmark EgoMask, but fine-tuning on EgoMask-Train yields significant improvements, while preserving performance on exocentric datasets. Our work thus provides essential resources and insights for advancing egocentric video understanding. Our code is available at https://github.com/LaVi-Lab/EgoMask . Shuo Liang, Yiwu Zhong, Zi-Yuan Hu, Yeyao Tao, Liwei Wang 0009 |
ICCV | 5 |
| 2025 | AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningabstractLarge language models (LLMs) have enabled the creation of multi-modal LLMs that exhibit strong comprehension of visual data such as images and videos. However, these models usually rely on extensive visual tokens from visual encoders, leading to high computational demands, which limits their applicability in resource-constrained environments and for long-context tasks. In this work, we propose a training-free adaptive inference method for multi-modal LLMs that can accommodate a broad range of efficiency requirements with a minimum performance drop. Our method consists of a) iterative token merging based on embedding similarity before LLMs, and b) progressive token pruning within LLM layers based on multi-modal importance. With a minimalist design, our method can be applied to both video and image LLMs. Extensive experiments on diverse video and image benchmarks demonstrate that our method substantially reduces computation load (e.g., a $\textbf{7-fold}$ reduction in FLOPs) while preserving the performance of video and image LLMs. Further, at a similar computational cost, our method outperforms the state-of-the-art methods in long video understanding (e.g., $\textbf{+4.6}$ on MLVU). Additionally, our in-depth analysis provides insights into token redundancy and LLM layer behaviors, offering guidance for future research in designing efficient multi-modal LLMs. Our code is available at https://github.com/LaVi-Lab/AIM. Yiwu Zhong, Zhuoming Liu 0001, Yin Li 0003, Liwei Wang 0009 |
ICCV | 4 |
| 2025 | Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsabstractPrevious research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point clouds or reconstructed Bird's-Eye View (BEV) maps. In our research, we advance this field by enhancing the capability of MLLMs to understand and reason in 3D spaces directly from video data, without the need for additional 3D input. We propose a novel and efficient method called the Video-3D Geometry Large Language Model (VG LLM). Our approach utilizes a 3D visual geometry encoder to extract 3D prior information from video sequences. This information is then integrated with visual tokens and input into the MLLM. Extensive experiments have shown that our method has achieved substantial improvements in various tasks related to 3D scene understanding and spatial reasoning, all directly learned from video sources. Impressively, our 4B model, which does not rely on explicit 3D data inputs, achieves competitive results compared to existing state-of-the-art methods, and even surpasses the Gemini-1.5-Pro in the VSI-Bench evaluations. Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang 0009 |
NeurIPS | 4 |
| 2025 | A Mutual Supervision Framework for Referring Expression Segmentation and Generation
Shijia Huang, Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Lei Zhang 0001, Liwei Wang 0009 |
Int. J. Comput. Vis. | 6 |
| 2024 | Making Long-Context Language Models Better Multi-Hop ReasonersabstractRecent advancements in long-context modeling have enhanced language models (LMs) for complex tasks across multiple NLP applications.Despite this progress, we find that these models struggle with multi-hop reasoning and exhibit decreased performance in the presence of noisy contexts.In this paper, we introduce Reasoning with Attributions, a novel approach that prompts LMs to supply attributions for each assertion during their reasoning.We validate our approach through experiments on three multi-hop datasets, employing both proprietary and open-source models, and demonstrate its efficacy and resilience.Furthermore, we explore methods to augment reasoning capabilities via fine-tuning and offer an attribution-annotated dataset and a specialized training strategy.Our fine-tuned model achieves competitive performance on multi-hop reasoning benchmarks, closely paralleling proprietary LMs such as ChatGPT and Claude-instant 1 . Yanyang Li, Shuo Liang, Michael R. Lyu, Liwei Wang 0009 |
ACL (1) | 4 |
| 2024 | Towards Learning a Generalist Model for Embodied NavigationabstractBuilding a generalist agent that can interact with the world is the intriguing target of AI systems, thus spurring the research for embodied navigation, where an agent is required to navigate according to instructions or respond to queries. Despite the major progress attained, previous works primarily focus on task-specific agents and lack gen-eralizability to unseen scenarios. Recently, LLMs have pre-sented remarkable capabilities across various fields, and provided a promising opportunity for embodied navigation. Drawing on this, we propose the first generalist model for embodied navigation, NaviLLM. It adapts LLMs to em-bodied navigation by introducing schema-based instruction. The schema-based instruction flexibly casts various tasks into generation problems, thereby unifying a wide range of tasks. This approach allows us to integrate di-verse data sources from various datasets into the training, equipping NaviLLM with a wide range of capabilities required by embodied navigation. We conduct extensive ex-periments to evaluate the performance and generalizability of our model. The experimental results demonstrate that our unified model achieves state-of-the-art performance on CVDN, SOON, and ScanQA. Specifically, it surpasses the previous stats-of-the-art method by a significant margin of 29% in goal progress on CVDN. Moreover, our model also demonstrates strong generalizability and presents im-pressive results on unseen tasks, e.g. embodied question answering and 3D captioning. Our code is available at https://github.com/LaVi-Lab/NaviLLM. Duo Zheng, Shijia Huang, Lin Zhao 0016, Yiwu Zhong, Liwei Wang 0009 |
CVPR | 5 |
| 2024 | Beyond Embeddings: The Promise of Visual Table in Visual ReasoningabstractVisual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations.Despite the success of CLIP-type visual embeddings, they often lack access to world knowledge critical for visual reasoning.In this work, we propose Visual Table, a novel form of visual representation tailored for visual reasoning.Visual tables are constructed as hierarchical descriptions of visual scenes, featuring a scene description and multiple object-centric descriptions covering categories, attributes, and knowledge.Thanks to the structural and textual formats, visual tables offer unique properties over mere visual embeddings, such as explainability and controllable editing.Furthermore, they deliver instance-level world knowledge and detailed attributes that are essential for visual reasoning.To create visual tables, we develop a generator trained on the dataset with collected, small-scale annotations.Extensive results on 11 visual reasoning benchmarks demonstrate that the generated visual tables significantly outperform previous structural and text-based representations.Moreover, they consistently enhance state-of-theart multi-modal large language models across diverse benchmarks, showcasing their potential for advancing visual reasoning tasks.Our code is available at https://github.com/ LaVi-Lab/Visual-Table. Yiwu Zhong, Zi-Yuan Hu, Michael R. Lyu, Liwei Wang 0009 |
EMNLP | 4 |
| 2023 | MVP-Tuning: Multi-View Knowledge Retrieval with Prompt Tuning for Commonsense ReasoningabstractYongfeng Huang, Yanyang Li, Yichong Xu, Lin Zhang, Ruyi Gan, Jiaxing Zhang, Liwei Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yongfeng Huang 0001, Yanyang Li, Yichong Xu, Ruyi Gan, Jiaxing Zhang 0001, Liwei Wang 0009 |
ACL (1) | 7 |
| 2023 | Learning Preference Model for LLMs via Automatic Preference Data GenerationabstractDespite the advanced capacities of the stateof-the-art large language models (LLMs), they suffer from issues of hallucination, stereotype, etc. Preference models play an important role in LLM alignment, yet training preference models predominantly rely on human-annotated data.This reliance limits their versatility and scalability.In this paper, we propose learning the preference model for LLMs via automatic preference data generation (AutoPM).Our approach involves both In-Breadth Data Generation, which elicits pairwise preference data from LLMs following the helpful-honestharmless (HHH) criteria, and In-Depth Data Generation, which enriches the dataset with responses spanning a wide quality range.With HHH-guided preference data, our approach simultaneously enables the LLMs to learn human preferences and align with human values.Quantitative assessments on five benchmark datasets demonstrate the reliability and potential of AutoPM, pointing out a more general and scalable way to improve LLM performance. Shijia Huang, Jianqiao Zhao, Yanyang Li, Liwei Wang 0009 |
EMNLP | 4 |
| 2023 | VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity ControlabstractAs the model size of pre-trained language models (PLMs) grows rapidly, full fine-tuning becomes prohibitively expensive for model training and storage. In vision-and-language (VL), parameter-efficient tuning (PET) techniques are proposed to integrate modular modifications (e.g., Adapter and LoRA) into encoder-decoder PLMs. By tuning a small set of trainable parameters, these techniques perform on par with full fine-tuning. However, excessive modular modifications and neglecting the functionality gap between the encoders and decoders can lead to performance degradation, while existing PET techniques (e.g., VL-Adapter) overlook these critical issues. In this paper, we propose a Vision-and-Language Parameter-Efficient Tuning (VL-PET) framework to impose effective control over modular modifications via a novel granularity-controlled mechanism. Considering different granularity-controlled matrices generated by this mechanism, a variety of model-agnostic VL-PET modules can be instantiated from our framework for better efficiency and effective-ness trade-offs. We further propose lightweight PET module designs to enhance VL alignment and modeling for the encoders and maintain text generation for the decoders. Extensive experiments conducted on four image-text tasks and four video-text tasks demonstrate the efficiency, effectiveness and transferability of our VL-PET framework. In particular, our VL-PETlargewith lightweight PET module designs significantly outperforms VL-Adapter by 2.92% (3.41%) and LoRA by 3.37% (7.03%) with BART-base (T5-base) on image-text tasks. Furthermore, we validate the enhanced effect of employing our VL-PET designs on existing PET techniques, enabling them to achieve significant performance improvements. Our code is available at https://github.com/HenryHZY/VL-PET. Zi-Yuan Hu, Yanyang Li, Michael R. Lyu, Liwei Wang 0009 |
ICCV | 4 |
| 2023 | Conditional Temporal Variational AutoEncoder for Action Video Prediction
Xiaogang Xu 0002, Yi Wang 0074, Liwei Wang 0009, Bei Yu 0001, Jiaya Jia |
Int. J. Comput. Vis. | 3 |
| 2023 | Fully Convolutional Networks for Panoptic Segmentation With Point-Based SupervisionabstractIn this paper, we present a conceptually simple, strong, and efficient framework for fully- and weakly-supervised panoptic segmentation, called Panoptic FCN. Our approach aims to represent and predict foreground things and background stuff in a unified fully convolutional pipeline, which can be optimized with point-based fully or weak supervision. In particular, Panoptic FCN encodes each object instance or stuff category with the proposed kernel generator and produces the prediction by convolving the high-resolution feature directly. With this approach, instance-aware and semantically consistent properties for things and stuff can be respectively satisfied in a simple generate-kernel-then-segment workflow. Without extra boxes for localization or instance separation, the proposed approach outperforms the previous box-based and -free models with high efficiency. Furthermore, we propose a new form of point-based annotation for weakly-supervised panoptic segmentation. It only needs several random points for both things and stuff, which dramatically reduces the annotation cost of human. The proposed Panoptic FCN is also proved to have much superior performance in this weakly-supervised setting, which achieves 82% of the fully-supervised performance with only 20 randomly annotated points per instance. Extensive experiments demonstrate the effectiveness and efficiency of Panoptic FCN on COCO, VOC 2012, Cityscapes, and Mapillary Vistas datasets. And it sets up a new leading benchmark for both fully- and weakly-supervised panoptic segmentation. Hengshuang Zhao, Xiaojuan Qi 0001, Yukang Chen, Lu Qi 0001, Liwei Wang 0009, Jian Sun 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and EfficiencyabstractStructured pruning has been extensively studied on monolingual pre-trained language models and is yet to be fully evaluated on their multilingual counterparts.This work investigates three aspects of structured pruning on multilingual pre-trained language models: settings, algorithms, and efficiency.Experiments on nine downstream tasks show several counterintuitive phenomena: for settings, individually pruning for each language does not induce a better result; for algorithms, the simplest method performs the best; for efficiency, a fast model does not imply that it is also small.To facilitate the comparison on all sparsity levels, we present Dynamic Sparsification, a simple approach that allows training the model once and adapting to different model sizes at inference.We hope this work fills the gap in the study of structured pruning on multilingual pre-trained models and sheds light on future research. Yanyang Li, Fuli Luo, Runxin Xu, Songfang Huang, Fei Huang 0002, Liwei Wang 0009 |
ACL (1) | 6 |
| 2022 | Multi-View Transformer for 3D Visual GroundingabstractThe 3D visual grounding task aims to ground a natural language description to the targeted object in a 3D scene, which is usually represented in 3D point clouds. Previous works studied visual grounding under specific views. The vision-language correspondence learned by this way can easily fail once the view changes. In this paper, we propose a Multi-View Transformer (MVT) for 3D visual grounding. We project the 3D scene to a multi-view space, in which the position information of the 3D scene under different views are modeled simultaneously and aggregated together. The multi-view space enables the network to learn a more robust multi-modal representation for 3D visual grounding and eliminates the dependence on specific views. Extensive experiments show that our approach significantly outperforms all state-of-the-art methods. Specifically, on Nr3D and Sr3D datasets, our method outperforms the best competitor by 11.2% and 7.1% and even surpasses recent work with extra 2D assistance by 5.9% and 6.6%. Our code is available at https://github.com/sega-hsj/MVT-3DVG. Shijia Huang, Jiaya Jia, Liwei Wang 0009 |
CVPR | 4 |
| 2022 | Stratified Transformer for 3D Point Cloud Segmentationabstract3D point cloud segmentation has made tremendous progress in recent years. Most current methods focus on aggregating local features, but fail to directly model long-range dependencies. In this paper, we propose Stratified Transformer that is able to capture long-range contexts and demonstrates strong generalization ability and high performance. Specifically, we first put forward a novel key sampling strategy. For each query point, we sample nearby points densely and distant points sparsely as its keys in a stratified way, which enables the model to enlarge the effective receptive field and enjoy long-range contexts at a low computational cost. Also, to combat the challenges posed by irregular point arrangements, we propose first-layer point embedding to aggregate local information, which facilitates convergence and boosts performance. Besides, we adopt contextual relative position encoding to adaptively capture position information. Finally, a memory-efficient implementation is introduced to overcome the issue of varying point numbers in each window. Extensive experiments demonstrate the effectiveness and superiority of our method on S3DIS, ScanNetv2 and ShapeNetPart datasets. Code is available at https://github.com/dvlab-research/Stratified-Transformer. Li Jiang 0009, Liwei Wang 0009, Hengshuang Zhao, Shu Liu 0005, Xiaojuan Qi 0001, Jiaya Jia |
CVPR | 4 |
| 2022 | Voxel Field Fusion for 3D Object DetectionabstractIn this work, we present a conceptually simple yet effective framework for cross-modality 3D object detection, named voxel field fusion. The proposed approach aims to maintain cross-modality consistency by representing and fusing augmented image features as a ray in the voxel field. To this end, the learnable sampler is first designed to sample vital features from the image plane that are projected to the voxel grid in a point-to-ray manner, which maintains the consistency in feature representation with spatial context. In addition, ray-wise fusion is conducted to fuse features with the supplemental context in the constructed voxel field. We further develop mixed augmentor to align feature-variant transformations, which bridges the modality gap in data augmentation. The proposed framework is demonstrated to achieve consistent gains in various bench-marks and outperforms previous fusion-based methods on KITTI and nuScenes datasets. Code is made available at https://github.com/dvlab-research/VFF11Part of the work was done in MEGVII Research.. Xiaojuan Qi 0001, Yukang Chen, Liwei Wang 0009, Jian Sun 0001, Jiaya Jia |
CVPR | 4 |
| 2022 | DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation
Zhuotao Tian, Xiaogang Xu 0002, Ying-Cong Chen, Shu Liu 0005, Hengshuang Zhao, Liwei Wang 0009, Jiaya Jia |
ECCV (33) | 7 |
| 2022 | Eliciting Knowledge from Large Pre-Trained Models for Unsupervised Knowledge-Grounded ConversationabstractRecent advances in large-scale pre-training provide large models with the potential to learn knowledge from the raw text.It is thus natural to ask whether it is possible to leverage these large models as knowledge bases for downstream tasks.In this work, we answer the aforementioned question in unsupervised knowledge-grounded conversation.We explore various methods that best elicit knowledge from large models.Our human study indicates that, though hallucinations exist, large models post the unique advantage of being able to output common sense and summarize facts that cannot be directly retrieved from the search engine.To better exploit such generated knowledge in dialogue generation, we treat the generated knowledge as a noisy knowledge source and propose the posterior-based reweighing as well as the noisy training strategy.Empirical results on two benchmarks show advantages over the state-of-the-art methods. Yanyang Li, Jianqiao Zhao, Michael R. Lyu, Liwei Wang 0009 |
EMNLP | 4 |
| 2022 | FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act FlowsabstractDespite recent progress in open-domain dialogue evaluation, how to develop automatic metrics remains an open problem.We explore the potential of dialogue evaluation featuring dialog act information, which was hardly explicitly modeled in previous methods.However, defined at the utterance level in general, dialog act is of coarse granularity, as an utterance can contain multiple segments possessing different functions.Hence, we propose segment act, an extension of dialog act from utterance level to segment level, and crowdsource a largescale dataset for it.To utilize segment act flows, sequences of segment acts, for evaluation, we develop the first consensus-based dialogue evaluation framework, FlowEval.This framework provides a reference-free approach for dialog evaluation by finding pseudo-references.Extensive experiments against strong baselines on three benchmark datasets demonstrate the effectiveness and other desirable characteristics of our FlowEval, pointing out a potential path for better dialogue evaluation. Jianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji, Dong Yu 0001, Michael R. Lyu, Liwei Wang 0009 |
EMNLP | 7 |
| 2021 | Self-Supervised 3D Mesh Reconstruction From Single ImagesabstractRecent single-view 3D reconstruction methods reconstruct object’s shape and texture from a single image with only 2D image-level annotation. However, without explicit 3D attribute-level supervision, it is still difficult to achieve satisfying reconstruction accuracy. In this paper, we propose a Self-supervised Mesh Reconstruction (SMR) approach to enhance 3D mesh attribute learning process. Our approach is motivated by observations that (1) 3D attributes from interpolation and prediction should be consistent, and (2) feature representation of landmarks from all images should be consistent. By only requiring silhouette mask annotation, our SMR can be trained in an end-to- end manner and generalizes to reconstruct natural objects of birds, cows, motorbikes, etc. Experiments demonstrate that our approach improves both 2D supervised and unsupervised 3D mesh reconstruction on multiple datasets. We also show that our model can be adapted to other image synthesis tasks, e.g., novel view generation, shape transfer, and texture transfer, with promising results. Our code is publicly available at https://github.com/Jia-Research-Lab. Tao Hu 0011, Liwei Wang 0009, Xiaogang Xu 0002, Shu Liu 0005, Jiaya Jia |
CVPR | 2 |
| 2021 | Semi-Supervised Semantic Segmentation With Directional Context-Aware ConsistencyabstractSemantic segmentation has made tremendous progress in recent years. However, satisfying performance highly depends on a large number of pixel-level annotations. Therefore, in this paper, we focus on the semi-supervised segmentation problem where only a small set of labeled data is provided with a much larger collection of totally unlabeled images. Nevertheless, due to the limited annotations, models may overly rely on the contexts available in the training data, which causes poor generalization to the scenes un-seen before. A preferred high-level representation should capture the contextual information while not losing self-awareness. Therefore, we propose to maintain the context-aware consistency between features of the same identity but with different contexts, making the representations robust to the varying environments. Moreover, we present the Directional Contrastive Loss (DC Loss) to accomplish the consistency in a pixel-to-pixel manner, only requiring the feature with lower quality to be aligned towards its counterpart. In addition, to avoid the false-negative samples and filter the uncertain positive samples, we put forward two sampling strategies. Extensive experiments show that our simple yet effective method surpasses current state-of-the-art methods by a large margin and also generalizes well with extra image-level annotations. Zhuotao Tian, Li Jiang 0009, Shu Liu 0005, Hengshuang Zhao, Liwei Wang 0009, Jiaya Jia |
CVPR | 6 |
| 2021 | Fully Convolutional Networks for Panoptic SegmentationabstractIn this paper, we present a conceptually simple, strong, and efficient framework for panoptic segmentation, called Panoptic FCN. Our approach aims to represent and predict foreground things and background stuff in a unified fully convolutional pipeline. In particular, Panoptic FCN encodes each object instance or stuff category into a specific kernel weight with the proposed kernel generator and produces the prediction by convolving the high-resolution feature directly. With this approach, instance-aware and semantically consistent prosperties for things and stuff can be respectively satisfied in a simple generate-kernel-then-segment workflow. Without extra boxes for localization or instance separation, the proposed approach outperforms previous box-based and -free models with high efficiency on COCO, Cityscapes, and Mapillary Vistas datasets with single scale input. Our code is made publicly available at https://github.com/Jia-Research-Lab/PanopticFCN.1 Hengshuang Zhao, Xiaojuan Qi 0001, Liwei Wang 0009, Jian Sun 0001, Jiaya Jia |
CVPR | 4 |
| 2021 | Improving Weakly Supervised Visual Grounding by Contrastive Knowledge DistillationabstractWeakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training time, and propose a contrastive learning framework that accounts for both region-phrase and image-sentence matching. Our core innovation is the learning of a region-phrase score function, based on which an image-sentence score function is further constructed. Importantly, our region-phrase score function is learned by distilling from soft matching scores between the detected object names and candidate phrases within an image-sentence pair, while the image-sentence score function is supervised by ground-truth image-sentence pairs. The design of such score functions removes the need of object detection at test time, thereby significantly reducing the inference cost. Without bells and whistles, our approach achieves state-of-the-art results on visual phrase grounding, surpassing previous methods that require expensive object detectors at test time. Liwei Wang 0009, Jing Huang 0014, Yin Li 0003, Kun Xu 0005, Zhengyuan Yang, Dong Yu 0001 |
CVPR | 1 |
| 2021 | RAST: Domain-Robust Dialogue Rewriting as Sequence TaggingabstractThe task of dialogue rewriting aims to reconstruct the latest dialogue utterance by copying the missing content from the dialogue context.Until now, the existing models for this task suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset.We address this robustness issue by proposing a novel sequence-taggingbased model so that the search space is significantly reduced, yet the core of this task is still well covered.As a common issue of most tagging models for text generation, the model's outputs may lack fluency.To alleviate this issue, we inject the loss signal from BLEU or GPT-2 under a REINFORCE framework.Experiments show huge improvements of our model over the current state-of-the-art systems when transferring to another dataset. Linfeng Song, Liwei Wang 0009, Kun Xu 0005, Zhaopeng Tu, Dong Yu 0001 |
EMNLP (1) | 3 |
| 2021 | Deep Structured Instance Graph for Distilling Object DetectorsabstractEffectively structuring deep knowledge plays a pivotal role in transfer from teacher to student, especially in semantic vision tasks. In this paper, we present a simple knowledge structure to exploit and encode information inside the detection system to facilitate detector knowledge distillation. Specifically, aiming at solving the feature imbalance problem while further excavating the missing relation inside semantic instances, we design a graph whose nodes correspond to instance proposal-level features and edges represent the relation between nodes. To further refine this graph, we design an adaptive background loss weight to reduce node noise and background samples mining to prune trivial edges. We transfer the entire graph as encoded knowledge representation from teacher to student, capturing local and global information simultaneously.We achieve new state-of-the-art results on the challenging COCO object detection task with diverse student-teacher pairs on both one- and two-stage detectors. We also experiment with instance segmentation to demonstrate robustness of our method. It is notable that distilled Faster R-CNN with ResNet18-FPN and ResNet50-FPN yields 38.68 and 41.82 Box AP respectively on the COCO benchmark, Faster R-CNN with ResNet101-FPN significantly achieves 43.38 AP, which outperforms ResNet152- FPN teacher about 0.7 AP. Code: https://github.com/dvlab-research/Dsig. Pengguang Chen, Shu Liu 0005, Liwei Wang 0009, Jiaya Jia |
ICCV | 4 |
| 2021 | Learnable Boundary Guided Adversarial TrainingabstractPrevious adversarial training raises model robustness under the compromise of accuracy on natural data. In this paper, we reduce natural accuracy degradation. We use the model logits from one clean model to guide learning of another one robust model, taking into consideration that logits from the well trained clean model embed the most discriminative features of natural data, e.g., generalizable classifier boundary. Our solution is to constrain logits from the robust model that takes adversarial examples as input and makes it similar to those from the clean model fed with corresponding natural data. It lets the robust model inherit the classifier boundary of the clean model. Moreover, we observe such boundary guidance can not only preserve high natural accuracy but also benefit model robustness, which gives new insights and facilitates progress for the adversarial community. Finally, extensive experiments on CIFAR-10, CIFAR-100, and Tiny ImageNet testify to the effectiveness of our method. We achieve new state-of-the-art robustness on CIFAR-100 without additional real or synthetic data with auto-attack benchmark1. Our code is available at https://github.com/dvlab-research/LBGAT. Jiequan Cui, Shu Liu 0005, Liwei Wang 0009, Jiaya Jia |
ICCV | 3 |
| 2021 | SAT: 2D Semantics Assisted Training for 3D Visual Groundingabstract3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region. Point clouds are sparse, noisy, and contain limited semantic information compared with 2D images. These inherent limitations make the 3D visual grounding problem more challenging. In this study, we propose 2D Semantics Assisted Training (SAT) that utilizes 2D image semantics in the training stage to ease point-cloud-language joint representation learning and assist 3D visual grounding. The main idea is to learn auxiliary alignments between rich, clean 2D object representations and the corresponding objects or mentioned entities in 3D scenes. SAT takes 2D object semantics, i.e., object label, image feature, and 2D geometric feature, as the extra input in training but does not require such inputs during inference. By effectively utilizing 2D semantics in training, our approach boosts the accuracy on the Nr3D dataset from 37.7% to 49.2%, which significantly surpasses the non-SAT baseline with the identical network architecture and inference input. Our approach outperforms the state of the art by large margins on multiple 3D visual grounding datasets, i.e., +10.4% absolute accuracy on Nr3D, +9.9% on Sr3D, and +5.6% on ScanRef. Zhengyuan Yang, Songyang Zhang 0004, Liwei Wang 0009, Jiebo Luo 0001 |
ICCV | 3 |
| 2020 | MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningabstractGenerating multi-sentence descriptions for videos is one of the most challenging captioning tasks due to its high requirements for not only visual relevance but also discoursebased coherence across the sentences in the paragraph.Towards this goal, we propose a new approach called Memory-Augmented Recurrent Transformer (MART), which uses a memory module to augment the transformer architecture.The memory module generates a highly summarized memory state from the video segments and the sentence history so as to help better prediction of the next sentence (w.r.t.coreference and repetition aspects), thus encouraging coherent paragraph generation.Extensive experiments, human evaluations, and qualitative analyses on two popular datasets ActivityNet Captions and YouCookII show that MART generates more coherent and less repetitive paragraph captions than baseline methods, while maintaining relevance to the input video events. 1 Jie Lei 0003, Liwei Wang 0009, Yelong Shen, Dong Yu 0001, Tamara L. Berg, Mohit Bansal |
ACL | 2 |
| 2020 | Improving One-Stage Visual Grounding by Recursive Sub-query Construction
Zhengyuan Yang, Liwei Wang 0009, Jiebo Luo 0001 |
ECCV (14) | 3 |
| 2020 | Comprehensive Image Captioning via Scene Graph Decomposition
Yiwu Zhong, Liwei Wang 0009, Jianshu Chen, Dong Yu 0001, Yin Li 0003 |
ECCV (14) | 2 |
| 2019 | Cross-lingual Knowledge Graph Alignment via Graph Matching Neural NetworkabstractPrevious cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we introduce the topic entity graph, a local sub-graph of an entity, to represent entities with their contextual information in KG. From this view, the KB-alignment task can be formulated as a graph matching problem; and we further propose a graph-attention based solution, which first matches all entities in two topic entity graphs, and then jointly model the local matching information to derive a graph-level matching vector. Experiments show that our model outperforms previous state-of-the-art methods by a large margin. Kun Xu 0005, Liwei Wang 0009, Mo Yu, Yansong Feng 0002, Yan Song 0003, Zhiguo Wang 0006, Dong Yu 0001 |
ACL (1) | 2 |
| 2019 | Fast, Diverse and Accurate Image Captioning Guided by Part-Of-SpeechabstractImage captioning is an ambiguous problem, with many suitable captions for an image. To address ambiguity, beam search is the de facto method for sampling multiple captions. However, beam search is computationally expensive and known to produce generic captions. To address this concern, some variational auto-encoder (VAE) and generative adversarial net (GAN) based methods have been proposed. Though diverse, GAN and VAE are less accurate. In this paper, we first predict a meaningful summary of the image, then generate the caption based on that summary. We use part-of-speech as summaries, since our summary should drive caption generation. We achieve the trifecta: (1) High accuracy for the diverse captions as evaluated by standard captioning metrics and user studies; (2) Faster computation of diverse captions compared to beam search and diverse beam search; and (3) High diversity as evaluated by counting novel sentences, distinct n-grams and mutual overlap (i.e., mBleu-4) scores. Aditya Deshpande, Jyoti Aneja, Liwei Wang 0009, Alexander G. Schwing, David A. Forsyth |
CVPR | 3 |
| 2019 | A Fast and Accurate One-Stage Approach to Visual GroundingabstractWe propose a simple, fast, and accurate one-stage approach to visual grounding, inspired by the following insight. The performances of existing propose-and-rank two-stage methods are capped by the quality of the region candidates they propose in the first stage - if none of the candidates could cover the ground truth region, there is no hope in the second stage to rank the right region to the top. To avoid this caveat, we propose a one-stage model that enables end-to-end joint optimization. The main idea is as straightforward as fusing a text query's embedding into the YOLOv3 object detector, augmented by spatial features so as to account for spatial mentions in the query. Despite being simple, this one-stage approach shows great potential in terms of both accuracy and speed for both phrase localization and referring expression comprehension, according to our experiments. Given these results along with careful investigations into some popular region proposals, we advocate for visual grounding a paradigm shift from the conventional two-stage methods to the one-stage framework. Zhengyuan Yang, Boqing Gong, Liwei Wang 0009, Wenbing Huang 0001, Dong Yu 0001, Jiebo Luo 0001 |
ICCV | 3 |
| 2019 | Learning Two-Branch Neural Networks for Image-Text Matching TasksabstractImage-language matching tasks have recently attracted a lot of attention in the computer vision field. These tasks include image-sentence matching, i.e., given an image query, retrieving relevant sentences and vice versa, and region-phrase matching or visual grounding, i.e., matching a phrase to relevant regions. This paper investigates two-branch neural networks for learning the similarity between these two data modalities. We propose two network structures that produce different output representations. The first one, referred to as an embedding network, learns an explicit shared latent embedding space with a maximum-margin ranking loss and novel neighborhood constraints. Compared to standard triplet sampling, we perform improved neighborhood sampling that takes neighborhood information into consideration while constructing mini-batches. The second network structure, referred to as a similarity network, fuses the two branches via element-wise product and is trained with regression loss to directly predict a similarity score. Extensive experiments show that our networks achieve high accuracies for phrase localization on the Flickr30K Entities dataset and for bi-directional image-sentence retrieval on Flickr30K and MSCOCO datasets. Liwei Wang 0009, Yin Li 0003, Jing Huang 0014, Svetlana Lazebnik |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Learning structural motif representations for efficient protein structure searchabstractMotivation: Given a protein of unknown function, fast identification of similar protein structures from the Protein Data Bank (PDB) is a critical step for inferring its biological function. Such structural neighbors can provide evolutionary insights into protein conformation, interfaces and binding sites that are not detectable from sequence similarity. However, the computational cost of performing pairwise structural alignment against all structures in PDB is prohibitively expensive. Alignment-free approaches have been introduced to enable fast but coarse comparisons by representing each protein as a vector of structure features or fingerprints and only computing similarity between vectors. As a notable example, FragBag represents each protein by a 'bag of fragments', which is a vector of frequencies of contiguous short backbone fragments from a predetermined library. Despite being efficient, the accuracy of FragBag is unsatisfactory because its backbone fragment library may not be optimally constructed and long-range interacting patterns are omitted. Results: Here we present a new approach to learning effective structural motif presentations using deep learning. We develop DeepFold, a deep convolutional neural network model to extract structural motif features of a protein structure. We demonstrate that DeepFold substantially outperforms FragBag on protein structural search on a non-redundant protein structure database and a set of newly released structures. Remarkably, DeepFold not only extracts meaningful backbone segments but also finds important long-range interacting motifs for structural comparison. We expect that DeepFold will provide new insights into the evolution and hierarchical organization of protein structural motifs. Availability and implementation: https://github.com/largelymfs/DeepFold. Yang Liu 0097, Liwei Wang 0009, Jian Peng 0001 |
Bioinform. | 3 |
| 2017 | Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding SpaceabstractThis paper explores image caption generation using conditional variational auto-encoders (CVAEs). Standard CVAEs with a fixed Gaussian prior yield descriptions with too little variability. Instead, we propose two models that explicitly structure the latent space around K components corresponding to different types of image content, and combine components to create priors for images that contain multiple types of content simultaneously (e.g., several kinds of objects). Our first model uses a Gaussian Mixture model (GMM) prior, while the second one defines a novel Additive Gaussian (AG) prior that linearly combines component means. We show that both models produce captions that are more diverse and more accurate than a strong LSTM baseline or a “vanilla” CVAE with a fixed Gaussian prior, with AG-CVAE showing particular promise. Liwei Wang 0009, Alexander G. Schwing, Svetlana Lazebnik |
NIPS | 1 |
| 2017 | Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang 0009, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik |
Int. J. Comput. Vis. | 2 |
| 2016 | Learning Deep Structure-Preserving Image-Text EmbeddingsabstractThis paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a largemargin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and textto-image retrieval. Our method achieves new state-of-theart results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset. Liwei Wang 0009, Yin Li 0003, Svetlana Lazebnik |
CVPR | 1 |
| 2015 | Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence ModelsabstractThe Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains linking mentions of the same entities in images, as well as 276k manually annotated bounding boxes corresponding to each entity. Such annotation is essential for continued progress in automatic image description and grounded language understanding. We present experiments demonstrating the usefulness of our annotations for text-to-image reference resolution, or the task of localizing textual entity mentions in an image, and for bidirectional image-sentence retrieval. These experiments confirm that we can further improve the accuracy of state-of-the-art retrieval methods by training with explicit region-to-phrase correspondence, but at the same time, they show that accurately inferring this correspondence given an image and a caption remains really challenging. Bryan A. Plummer, Liwei Wang 0009, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik |
ICCV | 2 |
| 2014 | Multi-scale Orderless Pooling of Deep Convolutional Activation Features
Yunchao Gong, Liwei Wang 0009, Svetlana Lazebnik |
ECCV (7) | 2 |
| 2014 | Improving Image-Sentence Embeddings Using Large Weakly Annotated Photo Collections
Yunchao Gong, Liwei Wang 0009, Micah Hodosh, Julia Hockenmaier, Svetlana Lazebnik |
ECCV (4) | 2 |
| 2014 | Learning to Predict from Crowdsourced Data
Wei Bi, Liwei Wang 0009, James T. Kwok, Zhuowen Tu |
UAI | 2 |
| 2012 | Discriminative Clustering via Generative Feature MappingabstractExisting clustering methods can be roughly classified into two categories: generative and discriminative approaches. Generative clustering aims to explain the data and thus is adaptive to the underlying data distribution; discriminative clustering, on the other hand, emphasizes on finding partition boundaries. In this paper, we take the advantages of both models by coupling the two paradigms through feature mapping derived from linearizing Bayesian classifiers. Such the feature mapping strategy maps nonlinear boundaries of generative clustering to linear ones in the feature space where we explicitly impose the maximum entropy principle. We also propose the unified probabilistic framework, enabling solvers using standard techniques. Experiments on a variety of datasets bear out the notable benefit of our method in terms of adaptiveness and robustness. Liwei Wang 0009, Zhuowen Tu, Jiaya Jia |
AAAI | 1 |
| 2012 | Learning sparse covariance patterns for natural scenesabstractFor scene classification, patch-level linear features do not always work as well as handcrafted features. In this paper, we present a new model to greatly improve the usefulness of linear features in classification by introducing co-variance patterns. We analyze their properties, discuss the fundamental importance, and present a generative model to properly utilize them. With this set of covariance information, in our framework, even the most naive linear features that originally lack the vital ability in classification become powerful. Experiments show that the performance of our new covariance model based on linear features is comparable with or even better than handcrafted features in scene classification. Liwei Wang 0009, Yin Li 0003, Jiaya Jia, Jian Sun 0001, David P. Wipf, James M. Rehg |
CVPR | 1 |
| 2012 | Bayesian Face Revisited: A Joint Formulation
Dong Chen 0003, Xudong Cao, Liwei Wang 0009, Fang Wen 0001, Jian Sun 0001 |
ECCV (3) | 3 |