EDBT 2026 Demo / reviewers in the wild / expert
Yi Wang 0074
dblp:17/221-74
· DBLP profile ↗
48ranked-venue papers
9as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 9 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured ReasoningabstractTourism and travel planning increasingly rely on digital assistance, yet existing multimodal AI systems often lack specialized knowledge and contextual understanding of urban environments. We present TraveLLaMA, a specialized multimodal language model designed for comprehensive travel assistance. Our work addresses the fundamental challenge of developing practical AI travel assistants through three key contributions: (1) TravelQA, a novel dataset of 265k question-answer pairs combining 160k text QA from authentic travel sources, 100k vision-language QA featuring maps and location imagery, and 5k expert-annotated Chain-of-Thought reasoning examples; (2) Travel-CoT, a structured reasoning framework that decomposes travel queries into spatial, temporal, and practical dimensions, improving answer accuracy by 10.8% while providing interpretable decision paths; and (3) an interactive agent system validated through extensive user studies. Through fine-tuning experiments on state-of-the-art vision-language models (LLaVA, Qwen-VL, Shikra), we achieve 6.2-9.4% base improvements, further enhanced by Travel-CoT reasoning. Our model demonstrates superior capabilities in contextual travel recommendations, map interpretation, and scene understanding while providing practical information such as operating hours and cultural insights. User studies with 500 participants show TraveLLaMA achieves a System Usability Scale score of 82.5, significantly outperforming general-purpose models and establishing new standards for multimodal travel assistance systems. Meng Chu, Yukang Chen, Haokun Gui, Shaozuo Yu, Yi Wang 0074, Jiaya Jia |
AAAI | 5 |
| 2026 | VideoChat-A1: Thinking with Long Videos by Chain-of-Shot ReasoningabstractRecent advances in video understanding have been driven by MLLMs. But these MLLMs are good at analyzing short videos, while suffering from difficulties in understanding videos with a longer context. To address this difficulty, several agent paradigms have recently been proposed, using MLLMs as agents for retrieving extra contextual knowledge in a long video. However, most existing agents ignore the key fact that a long video is composed with multiple shots, i.e., to answer the user question from a long video, it is critical to deeply understand its relevant shots like human. Without such insight, these agents often mistakenly find redundant even noisy temporal context, restricting their capacity for long video understanding. To fill this gap, we propose VideoChat-A1, a novel long video agent paradigm. Different from the previous works, our VideoChat-A1 can deeply think with long videos, via a distinct chain-of-shot reasoning paradigm. More specifically, it can progressively select the relevant shots of user question, and look into these shots in a coarse-to-fine partition. By multi-modal reasoning along the shot chain, VideoChat-A1 can effectively mimic step-by-step human thinking process, allowing the interactive discovery of preferable temporal context for thoughtful understanding in long videos. Extensive experiments show that, VideoChat-A1 achieves the state-of-the-art performance on the mainstream long video QA benchmarks, e.g., it achieves 77.0 on VideoMME(w/ subs) and 70.1 on EgoSchema, outperforming its strong baselines (e.g., InternVL2.5-8B and InternVideo2.5-8B), by up to 10.1% and 6.2%. Compared to leading closed-source GPT-4o and Gemini 1.5 Pro, VideoChat-A1 offers competitive accuracy, but only with 7% input frames and 12% inference time on average. Zikang Wang, Zhengrong Yue, Yi Wang 0074, Yu Qiao 0001, Limin Wang 0002, Yali Wang 0001 |
AAAI | 4 |
| 2026 | Causal Prompts for Open-Vocabulary Video Instance SegmentationabstractOpen-vocabulary Video Instance Segmentation addresses the challenging task of detecting, segmenting, and tracking objects in videos, including categories not encountered during training. However, existing approaches often overlook rich temporal cues from preceding frames, limiting their ability to leverage causal context for robust open-world generalization. To bridge this gap, we propose CPOVIS, a novel framework that introduces causal prompts-dynamically propagated visual and taxonomy prompts from historical frames-to enhance temporal reasoning and semantic consistency. Built upon a Mask2Former architecture with a CLIP backbone, CPOVIS integrates three core innovations: 1) PromptCLIP, which aligns cross-modal embeddings while preserving open-vocabulary capabilities; 2) a Visual Prompt Injector that propagates object-level features to maintain spatial-temporal coherence; and 3) a Taxonomy Prompt Infuser that leverages hierarchical semantic relationships to stabilize unseen category recognition. Furthermore, we introduce a contrastive learning strategy to disentangle object representations across frames and adapt the Segment Anything Model (SAM2) to boost open-vocabulary segmentation and tracking capacity in open-vocabulary video scenarios. Extensive experiments on seven challenging open- and closed-vocabulary video segmentation benchmarks demonstrate CPOVIS's state-of-the-art performance, outperforming existing methods by significant margins. Our findings highlight the critical role of causal prompt propagation in advancing video understanding in open-world scenarios. Rongkun Zheng, Lu Qi 0001, Xi Chen 0119, Yi Wang 0074, Kun Wang 0056, Yu Qiao 0001, Hengshuang Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task AlignmentabstractCurrent multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals although they give comprehensive perception and reasoning in a spectrum of vision applications. Recent studies either develop tool-using or unify specific visual tasks into the autoregressive framework, often at the expense of overall multimodal performance. To address this issue and enhance MLLMs with visual tasks in a scalable fashion, we propose Task Preference Optimization (TPO), a novel method that utilizes differentiable task preferences derived from typical fine-grained visual tasks. TPO introduces learnable task tokens that establish connections between multiple task-specific heads and the MLLM. By leveraging rich visual labels during training, TPO significantly enhances the MLLM’s multimodal capabilities and task-specific performance. Through multi-task co-training within TPO, we observe synergistic benefits that elevate individual task performance beyond what is achievable through single-task training methodologies. Our instantiation of this approach with VideoChat and LLaVA demonstrates an overall 14.6% improvement in multimodal performance compared to baseline models. Additionally, MLLM-TPO demonstrates robust zero-shot capabilities across various tasks, performing comparably to state-of-the-art supervised models. Ziang Yan, Yinan He, Chenting Wang, Kunchang Li 0002, Xinhao Li 0004, Xiangyu Zeng 0004, Zilei Wang, Yali Wang 0001, Yu Qiao 0001, Limin Wang 0002, Yi Wang 0074 |
CVPR | 12 |
| 2025 | Make Your Training Flexible: Towards Deployment-Efficient Video Models
Chenting Wang, Kunchang Li 0002, Tianxiang Jiang, Xiangyu Zeng 0004, Yi Wang 0074, Limin Wang 0002 |
ICCV | 5 |
| 2025 | VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosabstractWe present VRBench, the first long narrative video benchmark crafted for evaluating large models' multi-step reasoning capabilities, addressing limitations in existing evaluations that overlook temporal reasoning and procedural validity. It comprises 960 long videos (with an average duration of 1.6 hours), along with 8,243 human-labeled multi-step question-answering pairs and 25,106 reasoning steps with timestamps. These videos are curated via a multi-stage filtering process including expert inter-rater reviewing to prioritize plot coherence. We develop a human-AI collaborative framework that generates coherent reasoning chains, each requiring multiple temporally grounded steps, spanning seven types (e.g., event attribution, implicit inference). VRBench designs a multi-phase evaluation pipeline that assesses models at both the outcome and process levels. Apart from the MCQs for the final results, we propose a progress-level LLM-guided scoring metric to evaluate the quality of the reasoning chain from multiple dimensions comprehensively. Through extensive evaluations of 12 LLMs and 19 VLMs on VRBench, we undertake a thorough analysis and provide valuable insights that advance the field of multi-step reasoning. Jiashuo Yu, Yue Wu 0013, Meng Chu, Zhifei Ren, Zizheng Huang, Pei Chu, Yinan He, Zhenxiang Li, Zhongying Tu, Conghui He, Yu Qiao 0001, Yali Wang 0001, Yi Wang 0074, Limin Wang 0002 |
ICCV | 16 |
| 2025 | DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMsabstractIn video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input. While linear projectors are widely employed for encapsulation, they introduce semantic indistinctness and temporal incoherence when applied to videos. Conversely, the structure of resamplers shows promise in tackling these challenges, but an effective solution remains unexplored. Drawing inspiration from resampler structures, we introduce DisCo, a novel visual encapsulation method designed to yield semantically distinct and temporally coherent visual tokens for video MLLMs. DisCo integrates two key components: (1) A Visual Concept Discriminator (VCD) module, assigning unique semantics for visual tokens by associating them in pair with discriminative concepts in the video. (2) A Temporal Focus Calibrator (TFC) module, ensuring consistent temporal focus of visual tokens to video elements across every video frame. Through extensive experiments on multiple video MLLM frameworks, we demonstrate that DisCo remarkably outperforms previous state-of-the-art methods across a variety of video understanding benchmarks, while also achieving higher token efficiency thanks to the reduction of semantic indistinctness. The code: https://github.com/ZJHTerry18/DisCo. Jiahe Zhao, Rongkun Zheng, Yi Wang 0074, Helin Wang, Hengshuang Zhao |
ICCV | 3 |
| 2025 | ViLLa: Video Reasoning Segmentation with Large Language ModelabstractRecent efforts in video reasoning segmentation (VRS) integrate large language models (LLMs) with perception models to localize and track objects via textual instructions, achieving barely satisfactory results in simple scenarios. However, they struggled to discriminate and deduce the objects from user queries in more real-world scenes featured by long durations, multiple objects, rapid motion, and heavy occlusions. In this work, we analyze the underlying causes of these limitations, and present ViLLa: Video reasoning segmentation with Large Language Model. Remarkably, our ViLLa manages to tackle these challenges through multiple core innovations: (1) a context synthesizer that dynamically encodes the user intent with video contexts for accurate reasoning, resolving ambiguities in complex queries, and (2) a hierarchical temporal synchronizer that disentangles multi-object interactions across complex temporal scenarios by modelling multi-object interactions at local and global temporal scales. To enable efficient processing of long videos, ViLLa incorporates (3) a key segment sampler that adaptively partitions long videos into shorter but semantically dense segments for less redundancy. What's more, to promote research in this unexplored area, we construct a VRS benchmark, VideoReasonSeg, featuring different complex scenarios. Our model also exhibits impressive state-of-the-art results on VideoReasonSeg, Ref-YouTube-VOS, Ref-DAVIS17, MeViS, and ReVOS. Both quantitative and qualitative experiments demonstrate that our method effectively enhances video reasoning segmentation capabilities for multimodal LLMs. The code and dataset will be available at https://github.com/rkzheng99/ViLLa. Rongkun Zheng, Lu Qi 0001, Xi Chen 0072, Yi Wang 0074, Kun Wang 0056, Hengshuang Zhao |
ICCV | 4 |
| 2025 | Bootstrapping Language-Guided Navigation Learning with Self-Refining Data FlywheelabstractCreating high-quality data for training robust language-instructed agents is a long-lasting challenge in embodied AI. In this paper, we introduce a Self-Refining Data Flywheel (SRDF) that generates high-quality and large-scale navigational instruction-trajectory pairs by iteratively refining the data pool through the collaboration between two models, the instruction generator and the navigator, without any human-in-the-loop annotation.
Specifically, SRDF starts with using a base generator to create an initial data pool for training a base navigator, followed by applying the trained navigator to filter the data pool. This leads to higher-fidelity data to train a better generator, which can, in turn, produce higher-quality data for training the next-round navigator. Such a flywheel establishes a data self-refining process, yielding a continuously improved and highly effective dataset for large-scale language-guided navigation learning. Our experiments demonstrate that after several flywheel rounds, the navigator elevates the performance boundary from 70\% to 78\% SPL on the classic R2R test set, surpassing human performance (76\%) for the first time.
Meanwhile, this process results in a superior generator, evidenced by a SPICE increase from 23.5 to 26.2, better than all previous VLN instruction generation methods. Finally, we demonstrate the scalability of our method through increasing environment and instruction diversity, and
the generalization ability of our pre-trained navigator across various downstream navigation tasks, surpassing state-of-the-art methods by a large margin in all cases. Zun Wang 0001, Jialu Li 0001, Yicong Hong, Kunchang Li 0002, Shoubin Yu, Yi Wang 0074, Yu Qiao 0001, Yali Wang 0001, Mohit Bansal, Limin Wang 0002 |
ICLR | 7 |
| 2025 | TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningabstractMultimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for long video understanding, including a simple yet efficient framework to process long video sequence, a high-quality video dataset for grounded tuning of MLLMs, and a carefully-designed instruction tuning task to explicitly incorporate the grounding supervision in the traditional QA format. Specifically, based on VideoChat, we propose our long-video MLLM, coined as VideoChat-T, by implementing a token shuffling to compress long video tokens and introducing Temporal Adaptive Position Encoding (TAPE) to enhance the temporal awareness of visual representation. Meanwhile, we introduce the TimePro, a comprehensive grounding-centric instruction tuning dataset composed of 9 tasks and 349k high-quality grounded annotations. Notably, we design a new instruction tuning task type, called Temporal Grounded Caption, to peform detailed video descriptions with the corresponding time stamps prediction. This explicit temporal location prediction will guide MLLM to correctly attend on the visual content when generating description, and thus reduce the hallucination risk caused by the LLMs. Experimental results demonstrate that our TimeSuite provides a successful solution to enhance the long video understanding capability of short-form MLLM, achieving improvement of 5.6% and 6.8% on the benchmarks of Egoschema and VideoMME, respectively. In addition, VideoChat-T exhibits robust zero-shot temporal grounding capabilities, significantly outperforming the existing state-of-the-art MLLMs. After fine-tuning, it performs on par with the traditional supervised expert models. Xiangyu Zeng 0004, Kunchang Li 0002, Chenting Wang, Xinhao Li 0004, Tianxiang Jiang, Ziang Yan, Yansong Shi, Zhengrong Yue, Yi Wang 0074, Yali Wang 0001, Yu Qiao 0001, Limin Wang 0002 |
ICLR | 10 |
| 2025 | VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative PerceptionabstractInducing reasoning in multimodal large language models (MLLMs) is critical for achieving human-level perception and understanding. Existing methods mainly leverage LLM reasoning to analyze parsed visuals, often limited by static perception stages. This paper introduces Visual Test-Time Scaling (VTTS), a novel approach to enhance MLLMs' reasoning via iterative perception during inference. VTTS mimics humans' hierarchical attention by progressively refining focus on high-confidence spatio-temporal regions, guided by updated textual predictions. Specifically, VTTS employs an Iterative Perception (ITP) mechanism, incorporating reinforcement learning with spatio-temporal supervision to optimize reasoning. To support this paradigm, we also present VTTS-80K, a dataset tailored for iterative perception.
These designs allows a MLLM to enhance its performance by increasing its perceptual compute. Extensive experiments validate VTTS's effectiveness and generalization across diverse tasks and benchmarks. Our newly introduced Videochat-R1.5 model has achieved remarkable improvements, with an average increase of over 5\%, compared to robust baselines such as Qwen2.5VL-3B and -7B, across more than 15 benchmarks that encompass video conversation, video reasoning, and spatio-temporal perception. Ziang Yan, Yinan He, Xinhao Li 0004, Zhengrong Yue, Xiangyu Zeng 0004, Yali Wang 0001, Yu Qiao 0001, Limin Wang 0002, Yi Wang 0074 |
NeurIPS | 9 |
| 2025 | StreamForest: Efficient Online Video Understanding with Persistent Event MemoryabstractMultimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To address these challenges, we propose StreamForest, a novel architecture specifically designed for streaming video understanding. Central to StreamForest is the Persistent Event Memory Forest, a memory mechanism that adaptively organizes video frames into multiple event-level tree structures. This process is guided by penalty functions based on temporal distance, content similarity, and merge frequency, enabling efficient long-term memory retention under limited computational resources. To enhance real-time perception, we introduce a Fine-grained Spatiotemporal Window, which captures detailed short-term visual cues to improve current scene perception. Additionally, we present OnlineIT, an instruction-tuning dataset tailored for streaming video tasks. OnlineIT significantly boosts MLLM performance in both real-time perception and future prediction. To evaluate generalization in practical applications, we introduce ODV-Bench, a new benchmark focused on real-time streaming video understanding in autonomous driving scenarios. Experimental results demonstrate that StreamForest achieves the state-of-the-art performance, with accuracies of 77.3% on StreamingBench, 60.5% on OVBench, and 55.6% on OVO-Bench. In particular, even under extreme visual token compression (limited to 1024 tokens), the model retains 96.8% of its average accuracy in eight benchmarks relative to the default setting. These results underscore the robustness, efficiency, and generalizability of StreamForest for streaming video understanding. Xiangyu Zeng 0004, Kefan Qiu, Xinhao Li 0004, Ziang Yan, Xinhai Zhao, Yi Wang 0074, Limin Wang 0002 |
NeurIPS | 11 |
| 2025 | Seg-VAR: Image Segmentation with Visual Autoregressive ModelingabstractWhile visual autoregressive modeling (VAR) strategies have shed light on image generation with the autoregressive models, their potential for segmentation, a task that requires precise low-level spatial perception, remains unexplored. Inspired by the multi-scale modeling of classic Mask2Former-based models, we propose Seg-VAR, a novel framework that rethinks segmentation as a conditional autoregressive mask generation problem. This is achieved by replacing the discriminative learning with the latent learning process. Specifically, our method incorporates three core components: (1) an image encoder generating latent priors from input images, (2) a spatial-aware seglat (a latent expression of segmentation mask) encoder that maps segmentation masks into discrete latent tokens using a location-sensitive color mapping to distinguish instances, and (3) a decoder reconstructing masks from these latents. A multi-stage training strategy is introduced: first learning seglat representations via image-seglat joint training, then refining latent transformations, and finally aligning image-encoder-derived latents with seglat distributions. Experiments show Seg-VAR outperforms previous discriminative and generative methods on various segmentation tasks and validation benchmarks. By framing segmentation as a sequential hierarchical prediction task, Seg-VAR opens new avenues for integrating autoregressive reasoning into spatial-aware vision systems. Rongkun Zheng, Lu Qi 0001, Xi Chen 0072, Yi Wang 0074, Kun Wang 0056, Hengshuang Zhao |
NeurIPS | 4 |
| 2025 | VideoChat: chat-centric video understanding
Kunchang Li 0002, Yinan He, Yi Wang 0074, Yizhuo Li 0001, Wenhai Wang, Ping Luo 0002, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001 |
Sci. China Inf. Sci. | 3 |
| 2025 | LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models
Yaohui Wang 0001, Xin Ma 0031, Shangchen Zhou, Yi Wang 0074, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang 0001, Yuwei Guo 0002, Tianxing Wu 0002, Chenyang Si, Yuming Jiang 0003, Cunjian Chen, Chen Change Loy, Bo Dai 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002 |
Int. J. Comput. Vis. | 6 |
| 2025 | MSFM-UNET: enhancing medical image segmentation with multi-scale and multi-view frequency fusion
Qiang Gao 0017, Yi Wang 0074, Feiyan Zhou, Yong Li 0023, Bin Fang 0001, Lan Du 0002, Cunjian Chen |
Pattern Anal. Appl. | 2 |
| 2024 | MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkabstractWith the rapid development of Multi-modal Large language Models (MLLMs), a number of diagnostic bench-marks have recently emerged to evaluate the comprehension capabilities of these models. However, most bench-marks predominantly assess spatial understanding in the static image tasks, while overlooking temporal understanding in the dynamic video tasks. To alleviate this issue, we introduce a comprehensive Multi-modal Video understanding Benchmark, namely MVBench, which covers 20 chal-lenging video tasks that cannot be effectively solved with a single frame. Specifically, we first introduce a novel static-to-dynamic method to define these temporal-related tasks. By transforming various static tasks into dynamic ones, we enable the systematic generation of video tasks that require a broad spectrum of temporal skills, ranging from perception to cognition. Then, guided by the task definition, we au-tomatically convert public video annotations into multiple-choice QA to evaluate each task. On one hand, such a distinct paradigm allows us to build MVBench efficiently, without much manual intervention. On the other hand, it guarantees evaluation fairness with ground-truth video an-notations, avoiding the biased scoring of LLMs. More-over, we further develop a robust video MLLM baseline, i.e., VideoChat2, by progressive multi-modal training with di-verse instruction-tuning data. The extensive results on our MVBench reveal that, the existing MLLMs are far from sat-isfactory in temporal understanding, while our VideoChat2 largely surpasses these leading models by over 15% on MVBench. All models and data are available at https://github.com/OpenGVLab/Ask-Anything. Kunchang Li 0002, Yali Wang 0001, Yinan He, Yizhuo Li 0001, Yi Wang 0074, Yi Liu 0081, Zun Wang 0001, Jilan Xu, Guo Chen 0006, Ping Lou, Limin Wang 0002, Yu Qiao 0001 |
CVPR | 5 |
| 2024 | VideoMamba: State Space Model for Efficient Video Understanding
Kunchang Li 0002, Xinhao Li 0004, Yi Wang 0074, Yinan He, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001 |
ECCV (26) | 3 |
| 2024 | InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Yi Wang 0074, Kunchang Li 0002, Xinhao Li 0004, Jiashuo Yu, Yinan He, Guo Chen 0006, Baoqi Pei, Rongkun Zheng, Zun Wang 0001, Yansong Shi, Tianxiang Jiang, Jilan Xu, Hongjie Zhang 0002, Yifei Huang 0002, Yu Qiao 0001, Yali Wang 0001, Limin Wang 0002 |
ECCV (85) | 1 |
| 2024 | OSFENet: Object Spatiotemporal Feature Enhanced Network for Surgical Phase Recognition
Pingjie You, Hengqi Hu, Yi Wang 0074, Bin Fang 0001 |
ICIC (12) | 4 |
| 2024 | InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationabstractThis paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accompanied by detailed descriptions of total 4.1B words. Our core contribution is to develop a scalable approach to autonomously build a high-quality video-text dataset with large language models (LLM), thereby showcasing its efficacy in learning video-language representation at scale. Specifically, we utilize a multi-scale approach to generate video-related descriptions. Furthermore, we introduce ViCLIP, a video-text representation learning model based on ViT-L. Learned on InternVid via contrastive learning, this model demonstrates leading zero-shot action recognition and competitive video retrieval performance. Beyond basic video understanding tasks like recognition and retrieval, our dataset and model have broad applications. They are particularly beneficial for generating interleaved video-text data for learning a video-centric dialogue system, advancing video-to-text and text-to-video generation research. These proposed resources provide a tool for researchers and practitioners interested in multimodal video understanding and generation. Yi Wang 0074, Yinan He, Yizhuo Li 0001, Kunchang Li 0002, Jiashuo Yu, Xin Ma 0031, Xinhao Li 0004, Guo Chen 0006, Yaohui Wang 0001, Ping Luo 0002, Ziwei Liu 0002, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001 |
ICLR | 1 |
| 2024 | Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?abstractVideo understanding relies on accurate action detection for temporal analysis. However, existing mainstream methods have limitations in real-world applications due to their offline and closed-set evaluation approaches, as well as their dependence on manual annotations. To address these challenges and enable real-time action understanding in open-world scenarios, we propose OV-OAD, a zero-shot online action detector that leverages vision-language models and learns solely from text supervision. By introducing an object-centered decoder unit into a Transformer-based model, we aggregate frames with similar semantics using video-text correspondence. Extensive experiments on four action detection benchmarks demonstrate that OV-OAD outperforms other advanced zero-shot methods. Specifically, it achieves 37.5\% mean average precision on THUMOS’14 and 73.8\% calibrated average precision on TVSeries. This research establishes a robust baseline for zero-shot transfer in online action detection, enabling scalable solutions for open-world temporal understanding. The code will be available for download at \url{https://github.com/OpenGVLab/OV-OAD}. Yi Wang 0074, Jilan Xu, Yinan He, Zifan Song, Limin Wang 0002, Yu Qiao 0001, Cairong Zhao |
NeurIPS | 2 |
| 2024 | SyncVIS: Synchronized Video Instance SegmentationabstractRecent DETR-based methods have advanced the development of Video Instance Segmentation (VIS) through transformers' efficiency and capability in modeling spatial and temporal information. Despite harvesting remarkable progress, existing works follow asynchronous designs, which model video sequences via either video-level queries only or adopting query-sensitive cascade structures, resulting in difficulties when handling complex and challenging video scenarios. In this work, we analyze the cause of this phenomenon and the limitations of the current solutions, and propose to conduct synchronized modeling via a new framework named SyncVIS. Specifically, SyncVIS explicitly introduces video-level query embeddings and designs two key modules to synchronize video-level query with frame-level query embeddings: a synchronized video-frame modeling paradigm and a synchronized embedding optimization strategy. The former attempts to promote the mutual learning of frame- and video-level embeddings with each other and the latter divides large video sequences into small clips for easier optimization. Extensive experimental evaluations are conducted on the challenging YouTube-VIS 2019 & 2021 & 2022, and OVIS benchmarks, and SyncVIS achieves state-of-the-art results, which demonstrates the effectiveness and generality of the proposed approach. The code is available at https://github.com/rkzheng99/SyncVIS. Rongkun Zheng, Lu Qi 0001, Xi Chen 0072, Yi Wang 0074, Kun Wang 0056, Yu Qiao 0001, Hengshuang Zhao |
NeurIPS | 4 |
| 2023 | Advancing Pancreas Segmentation through the Patch-Adjust Fusion FrameworkabstractPancreas segmentation is pivotal for the timely diagnosis of pancreatic diseases, but it remains challenging due to organ’s small size, large spatial variations, and unclear boundaries. To enhance segmentation accuracy, researchers have explored various multi-modal coarse-to-fine (MMCF) methods, integrating distinct data-related factors such as stage, view, scale, and patch size. However, these methods often ask for multiple trained models, leading to significant storage and computational expenses. Moreover, these methods frequently treat different data-related factors separately, overlooking their potential relationships and constraining the feature representation capability of convolutional neural networks. To tackle these issues, we propose a data-efficient framework known as Patch-Adjust Fusion (PAF). PAF framework addresses these factors concurrently through dynamic patch flow adjustments, seamlessly integrating diverse elements while accounting for their interactions and dependencies. Consequently, the PAF framework demonstrates superior segmentation performance, diminished storage requisites, and expedited inference speeds compared to the MMCF framework. This is exemplified across two public CT pancreas datasets (NIH and MSD) and a private MR pancreas dataset (MR300). These results underscore its suitability for clinical applications, positioning the PAF framework as a promising solution for accurate and efficient pancreas segmentation in medical settings. Yi Wang 0074, Bin Fang 0001 |
BIBM | 2 |
| 2023 | VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingabstractScale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video foundation models with billions of parameters. This paper shows that video masked autoencoder (VideoMAE) is a scalable and general self-supervised pre-trainer for building video foundation models. We scale the VideoMAE in both model and data with a core design. Specifically, we present a dual masking strategy for efficient pre-training, with an encoder operating on a subset of video tokens and a decoder processing another subset of video tokens. Although VideoMAE is very efficient due to high masking ratio in encoder, masking decoder can still further reduce the overall computational cost. This enables the efficient pre-training of billion-level models in video. We also use a progressive training paradigm that involves an initial pre-training on a diverse multi-sourced unlabeled dataset, followed by a post-pre-training on a mixed labeled dataset. Finally, we successfully train a video ViT model with a billion parameters, which achieves a new state-of-the-art performance on the datasets of Kinetics (90.0% on K400 and 89.9% on K600) and Something-Something (68.7% on V1 and 77.0% on V2). In addition, we extensively verify the pre-trained video ViT models on a variety of downstream tasks, demonstrating its effectiveness as a general video representation learner. Limin Wang 0002, Bingkun Huang, Zhan Tong, Yinan He, Yi Wang 0074, Yali Wang 0001, Yu Qiao 0001 |
CVPR | 6 |
| 2023 | Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionabstractThis paper considers the problem of open-vocabulary semantic segmentation (OVS), that aims to segment objects of arbitrary classes beyond a pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled imagetext pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slotattention based binding module, then aligns the group tokens to corresponding caption embeddings. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, encouraging the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on four benchmark datasets, PASCAL VOC, PASCAL Context, COCO Object, and ADE20K. OVSegmentor achieves superior results over state-of-the-art approaches on PASCAL VOC using only 3% data (4M vs 134M) for pre-training. Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Yi Wang 0074, Yu Qiao 0001, Weidi Xie |
CVPR | 5 |
| 2023 | NeRFLiX: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-viewpoint MiXerabstractNeural radiance fields (NeRF) show great success in novel view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation inaccuracy. Even with high-quality training frames, the synthetic novel views produced by NeRF models still suffer from notable rendering artifacts, such as noise, blur, etc. Towards to improve the synthesis quality of NeRF-based approaches, we propose NeRFLiX, a general NeRF-agnostic restorer paradigm by learning a degradation-driven inter-viewpoint mixer. Specially, we design a NeRF-style degradation modeling approach and construct large-scale training data, enabling the possibility of effectively removing NeRF-native rendering artifacts for existing deep neural networks. Moreover, beyond the degradation removal, we propose an inter-viewpoint aggregation framework that is able to fuse highly related high-quality training images, pushing the performance of cutting-edge NeRF models to entirely new levels and producing highly photo-realistic synthetic views. Kun Zhou 0001, Wenbo Li 0001, Yi Wang 0074, Tao Hu 0011, Nianjuan Jiang, Xiaoguang Han 0001, Jiangbo Lu |
CVPR | 3 |
| 2023 | Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsabstractVideo Foundation Models (VFMs) have received limited exploration due to high computational costs and data scarcity. Previous VFMs rely on Image Foundation Models (IFMs), which face challenges in transferring to the video domain. Although VideoMAE has trained a robust ViT from limited data, its low-level reconstruction poses convergence difficulties and conflicts with high-level cross-modal alignment. This paper proposes a training-efficient method for temporal-sensitive VFMs that integrates the benefits of existing methods. To increase data efficiency, we mask out most of the low-semantics video tokens, but selectively align the unmasked tokens with IFM, which serves as the UnMasked Teacher (UMT). By providing semantic guidance, our method enables faster convergence and multi-modal friendliness. With a progressive pre-training framework, our model can handle various tasks including scene-related, temporal-related, and complex video-language understanding. Using only public sources for pre-training in 6 days on 32 A100 GPUs, our scratch-built ViT-L/16 achieves state-of-the-art performances on various video tasks. Kunchang Li 0002, Yali Wang 0001, Yizhuo Li 0001, Yi Wang 0074, Yinan He, Limin Wang 0002, Yu Qiao 0001 |
ICCV | 4 |
| 2023 | UniFormerV2: Unlocking the Potential of Image ViTs for Video UnderstandingabstractThe prolific performances of Vision Transformers (ViTs) in image tasks have prompted research into adapting the image ViTs for video tasks. However, the substantial gap between image and video impedes the spatiotemporal learning of these image-pretrained models. Though video-specialized models like UniFormer can transfer to the video domain more seamlessly, their unique architectures require prolonged image pretraining, limiting the scalability. Given the emergence of powerful open-source image ViTs, we propose unlocking their potential for video understanding with efficient UniFormer designs. We call the resulting model UniFormerV2, since it inherits the concise style of the Uni-Former block, while redesigning local and global relation aggregators that seamlessly integrate advantages from both ViTs and UniFormer. Our UniFormerV2 achieves state-of-the-art performances on 8 popular video benchmarks, including scene-related Kinetics-400/600/700, heterogeneous Moments in Time, temporal-related Something-Something V1/V2, and untrimmed ActivityNet and HACS. It is note-worthy that to the best of our knowledge, UniFormerV2 is the first to elicit 90% top-1 accuracy on Kinetics-400. Kunchang Li 0002, Yali Wang 0001, Yinan He, Yizhuo Li 0001, Yi Wang 0074, Limin Wang 0002, Yu Qiao 0001 |
ICCV | 5 |
| 2023 | Scaling Data Generation in Vision-and-Language NavigationabstractRecent research in language-guided visual navigation has demonstrated a significant demand for the diversity of traversable environments and the quantity of supervision for training generalizable agents. To tackle the common data scarcity issue in existing vision-and-language navigation datasets, we propose an effective paradigm for generating large-scale data for learning, which applies 1200+ photo-realistic environments from HM3D and Gibson datasets and synthesizes 4.9 million instruction-trajectory pairs using fully-accessible resources on the web. Importantly, we investigate the influence of each component in this paradigm on the agent’s performance and study how to adequately apply the augmented data to pre-train and fine-tune an agent. Thanks to our large-scale dataset, the performance of an existing agent can be pushed up (+11% absolute with regard to previous SoTA) to a significantly new best of 80% single-run success rate on the R2R test split by simple imitation learning. The long-lasting generalization gap between navigating in seen and unseen environments is also reduced to less than 1% (versus 8% in the previous best method). Moreover, our paradigm also facilitates different models to achieve new state-of-the-art navigation results on CVDN, REVERIE, and R2R in continuous environments. Zun Wang 0001, Jialu Li 0001, Yicong Hong, Yi Wang 0074, Qi Wu 0001, Mohit Bansal, Stephen Gould, Hao Tan 0002, Yu Qiao 0001 |
ICCV | 4 |
| 2023 | JourneyDB: A Benchmark for Generative Image UnderstandingabstractWhile recent advancements in vision-language models have had a transformative impact on multi-modal comprehension, the extent to which these models possess the ability to comprehend generated images remains uncertain. Synthetic images, in comparison to real data, encompass a higher level of diversity in terms of both content and style, thereby presenting significant challenges for the models to fully grasp. In light of this challenge, we introduce a comprehensive dataset, referred to as JourneyDB, that caters to the domain of generative images within the context of multi-modal visual understanding. Our meticulously curated dataset comprises 4 million distinct and high-quality generated images, each paired with the corresponding text prompts that were employed in their creation. Furthermore, we additionally introduce an external subset with results of another 22 text-to-image generative models, which makes JourneyDB a comprehensive benchmark for evaluating the comprehension of generated images. On our dataset, we have devised four benchmarks to assess the performance of generated image comprehension in relation to both content and style interpretation. These benchmarks encompass prompt inversion, style retrieval, image captioning, and visual question answering. Lastly, we evaluate the performance of state-of-the-art multi-modal models when applied to the JourneyDB dataset, providing a comprehensive analysis of their strengths and limitations in comprehending generated content. We anticipate that the proposed dataset and benchmarks will facilitate further research in the field of generative content understanding. The dataset is publicly available at https://journeydb.github.io. Keqiang Sun, Junting Pan, Yuying Ge, Hao Li 0069, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang 0074, Jifeng Dai, Yu Qiao 0001, Limin Wang 0002, Hongsheng Li 0001 |
NeurIPS | 10 |
| 2023 | TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance SegmentationabstractTraining on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific datasets, thus, it is appealing to jointly train models across the aggregation of datasets to enhance data volume and diversity. However, due to the heterogeneity in category space, as mask precision increase with the data volume, simply utilizing multiple datasets will dilute the attention of models on different taxonomy. Thus, increasing the data scale and enriching taxonomy space while improving classification precision is important. In this work, we analyze that providing extra taxonomy information can help models concentrate on specific taxonomy, and propose our model named Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation (TMT-VIS) to address this vital challenge. Specifically, we design a two-stage taxonomy aggregation module that first compiles taxonomy information from input videos and then aggregates these taxonomy priors into instance queries before the transformer decoder. We conduct extensive experimental evaluations on four popular and challenging benchmarks, including YouTube-VIS 2019, YouTube-VIS 2021, OVIS, and UVO. Our model shows significant improvement over the baseline solutions, and sets new state-of-the-art records on all these benchmarks. These appealing and encouraging results demonstrate the effectiveness and generality of our proposed approach. The code and trained models will be publicly available. Rongkun Zheng, Lu Qi 0001, Xi Chen 0072, Yi Wang 0074, Kun Wang 0056, Yu Qiao 0001, Hengshuang Zhao |
NeurIPS | 4 |
| 2023 | Conditional Temporal Variational AutoEncoder for Action Video Prediction
Xiaogang Xu 0002, Yi Wang 0074, Liwei Wang 0009, Bei Yu 0001, Jiaya Jia |
Int. J. Comput. Vis. | 2 |
| 2023 | Open World Entity SegmentationabstractWe introduce a new image segmentation task, called Entity Segmentation (ES), which aims to segment all visual entities (objects and stuffs) in an image without predicting their semantic labels. By removing the need of class label prediction, the models trained for such task can focus more on improving segmentation quality. It has many practical applications such as image manipulation and editing where the quality of segmentation masks is crucial but class labels are less important. We conduct the first-ever study to investigate the feasibility of convolutional center-based representation to segment things and stuffs in a unified manner, and show that such representation fits exceptionally well in the context of ES. More specifically, we propose a CondInst-like fully-convolutional architecture with two novel modules specifically designed to exploit the class-agnostic and non-overlapping requirements of ES. Experiments show that the models designed and trained for ES significantly outperforms popular class-specific panoptic segmentation models in terms of segmentation quality. Moreover, an ES model can be easily trained on a combination of multiple datasets without the need to resolve label conflicts in dataset merging, and the model trained for ES on one or more datasets can generalize very well to other test datasets of unseen domains. The code has been released at https://github.com/dvlab-research/Entity. Lu Qi 0001, Jason Kuen, Yi Wang 0074, Jiuxiang Gu, Hengshuang Zhao, Philip Torr 0001, Zhe Lin 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | MAT: Mask-Aware Transformer for Large Hole Image InpaintingabstractRecent studies have shown the importance of modeling long-range interactions in the inpainting problem. To achieve this goal, existing approaches exploit either standalone attention techniques or transformers, but usually under a low resolution in consideration of computational cost. In this paper, we present a novel transformer-based model for large hole inpainting, which unifies the merits of transformers and convolutions to efficiently process high-resolution images. We carefully design each component of our framework to guarantee the high fidelity and diversity of recovered images. Specifically, we customize an inpainting-oriented transformer block, where the attention module aggregates non-local information only from partial valid tokens, indicated by a dynamic mask. Extensive experiments demonstrate the state-of-the-art performance of the new model on multiple benchmark datasets. Code is released at https://github.com/fenglinglwb/MAT. Wenbo Li 0002, Zhe Lin 0001, Kun Zhou 0001, Lu Qi 0001, Yi Wang 0074, Jiaya Jia |
CVPR | 5 |
| 2022 | Towards Implicit Text-Guided 3D Shape GenerationabstractIn this work, we explore the challenging task of generating 3D shapes from text. Beyond the existing works, we propose a new approach for text-guided 3D shape generation, capable of producing high-fidelity shapes with colors that match the given text description. This work has several technical contributions. First, we decouple the shape and color predictions for learning features in both texts and shapes, and propose the word-level spatial transformer to correlate word features from text with spatial features from shape. Also, we design a cyclic loss to encourage consistency between text and shape, and introduce the shape IMLE to diversify the generated shapes. Further, we extend the framework to enable text-guided shape manipulation. Extensive experiments on the largest existing text-shape benchmark [10] manifest the superiority of this work. The code and the models are available at https://github.com/liuzhengzhe/Towards-Implicit-Text-Guided-Shape-Generation. Zhengzhe Liu, Yi Wang 0074, Xiaojuan Qi 0001, Chi-Wing Fu |
CVPR | 2 |
| 2022 | PalGAN: Image Colorization with Palette Generative Adversarial Networks
Yi Wang 0074, Menghan Xia, Lu Qi 0001, Yu Qiao 0001 |
ECCV (15) | 1 |
| 2022 | PointINS: Point-Based Instance SegmentationabstractIn this paper, we explore the mask representation in instance segmentation with Point-of-Interest (PoI) features. Differentiating multiple potential instances within a single PoI feature is challenging, because learning a high-dimensional mask feature for each instance using vanilla convolution demands a heavy computing burden. To address this challenge, we propose an instance-aware convolution. It decomposes this mask representation learning task into two tractable modules as instance-aware weights and instance-agnostic features. The former is to parametrize convolution for producing mask features corresponding to different instances, improving mask learning efficiency by avoiding employing several independent convolutions. Meanwhile, the latter serves as mask templates in a single point. Together, instance-aware mask features are computed by convolving the template with dynamic weights, used for the mask prediction. Along with instance-aware convolution, we propose PointINS, a simple and practical instance segmentation approach, building upon dense one-stage detectors. Through extensive experiments, we evaluated the effectiveness of our framework built upon RetinaNet and FCOS. PointINS in ResNet101 backbone achieves a 38.3 mask mean average precision (mAP) on COCO dataset, outperforming existing point-based methods by a large margin. It gives a comparable performance to the region-based Mask R-CNN K. He, G. Gkioxari, P. Dollár, and R. Girshick, "Mask R-CNN," in Proc. IEEE Int. Conf. Comput. Vis., 2017, pp. 2980-2988 with faster inference. Lu Qi 0001, Yi Wang 0074, Yukang Chen, Ying-Cong Chen, Xiangyu Zhang 0005, Jian Sun 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Multi-Scale Aligned Distillation for Low-Resolution DetectionabstractIn instance-level detection tasks (e.g., object detection), reducing input resolution is an easy option to improve runtime efficiency. However, this option traditionally hurts the detection performance much. This paper focuses on boosting performance of low-resolution models by distilling knowledge from a high- or multi-resolution model. We first identify the challenge of applying knowledge distillation (KD) to teacher and student networks that act on different input resolutions. To tackle it, we explore the idea of spatially aligning feature maps between models of varying input resolutions by shifting feature pyramid position and introduce aligned multi-scale training to train a multi-scale teacher that can distill its knowledge to a low-resolution student. Further, we propose crossing feature-level fusion to dynamically fuse teacher’s multi-resolution features to guide the student better. On several instance-level detection tasks and datasets, the low-resolution models trained via our approach perform competitively with high-resolution models trained via conventional multi-scale training, while outperforming the latter’s low-resolution models by 2.1% to 3.6% in terms of mAP. Our code is made publicly available at https://github.com/Jia-Research-Lab/MSAD. Lu Qi 0001, Jason Kuen, Jiuxiang Gu, Zhe Lin 0001, Yi Wang 0074, Yukang Chen, Jiaya Jia |
CVPR | 5 |
| 2021 | Image Synthesis via Semantic CompositionabstractIn this paper, we present a novel approach to synthesize realistic images based on their semantic layouts. It hypothesizes that for objects with similar appearance, they share similar representation. Our method establishes dependencies between regions according to their appearance correlation, yielding both spatially variant and associated representations. Conditioning on these features, we propose a dynamic weighted network constructed by spatially conditional computation (with both convolution and normalization). More than preserving semantic distinctions, the given dynamic network strengthens semantic relevance, benefiting global structure and detail synthesis. We demonstrate that our method gives the compelling generation performance qualitatively and quantitatively with extensive experiments on benchmarks. Yi Wang 0074, Lu Qi 0001, Ying-Cong Chen, Xiangyu Zhang 0005, Jiaya Jia |
ICCV | 1 |
| 2021 | Blood Vessel Segmentation Based on the 3D Residual U-NetabstractIn this paper, we propose blood vessel segmentation based on the 3D residual U-Net method. First, we integrate the residual block structure into the 3D U-Net. By exploring the influence of adding residual blocks at different positions in the 3D U-Net, we establish a novel and effective 3D residual U-Net. In addition, to address the challenges of pixel imbalance, vessel boundary segmentation, and small vessel segmentation, we develop a new weighted Dice loss function with a better effect than the weighted cross-entropy loss function. When training the model, we adopted a two-stage method from coarse-to-fine. In the fine stage, a local segmentation method of 3D sliding window is added. In the model testing phase, we used the 3D fixed-point method. Furthermore, we employ the 3D morphological closed operation to smooth the surfaces of vessels and volume analysis to remove noise blocks. To verify the accuracy and stability of our method, we compare our method with FCN, 3D DenseNet, and 3D U-Net. The experimental results indicate that our method has higher accuracy and better stability than the other studied methods and that the average Dice coefficients for hepatic veins and portal veins reach 71.7% and 76.5% in the coarse stage and 72.5% and 77.2% in the fine stage, respectively. In order to verify the robustness of the model, we conducted the same comparative experiment on the brain vessel datasets, and the average Dice coefficient reached 87.2%. Mulin Xin, Yi Wang 0074, Bin Fang 0001, Yongmei Xu, Chunhong Linghu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2020 | Attentive Normalization for Conditional Image GenerationabstractTraditional convolution-based generative adversarial networks synthesize images based on hierarchical local operations, where long-range dependency relation is implicitly modeled with a Markov chain. It is still not sufficient for categories with complicated structures. In this paper, we characterize long-range dependence with attentive normalization (AN), which is an extension to traditional instance normalization. Specifically, the input feature map is softly divided into several regions based on its internal semantic similarity, which are respectively normalized. It enhances consistency between distant regions with semantic correspondence. Compared with self-attention GAN, our attentive normalization does not need to measure the correlation of all locations, and thus can be directly applied to large-size feature maps without much computational burden. Extensive experiments on class-conditional image generation and semantic inpainting verify the efficacy of our proposed module. Yi Wang 0074, Ying-Cong Chen, Xiangyu Zhang 0005, Jian Sun 0001, Jiaya Jia |
CVPR | 1 |
| 2020 | VCNet: A Robust Approach to Blind Image Inpainting
Yi Wang 0074, Ying-Cong Chen, Xin Tao 0001, Jiaya Jia |
ECCV (25) | 1 |
| 2019 | Wide-Context Semantic Image ExtrapolationabstractThis paper studies the fundamental problem of extrapolating visual context using deep generative models, i.e., extending image borders with plausible structure and details. This seemingly easy task actually faces many crucial technical challenges and has its unique properties. The two major issues are size expansion and one-side constraints. We propose a semantic regeneration network with several special contributions and use multiple spatial related losses to address these issues. Our results contain consistent structures and high-quality textures. Extensive experiments are conducted on various possible alternatives and related methods. We also explore the potential of our method for various interesting applications that can benefit research in a variety of fields. Yi Wang 0074, Xin Tao 0001, Xiaoyong Shen, Jiaya Jia |
CVPR | 1 |
| 2019 | Liver Vessels Segmentation Based on 3d Residual U-NETabstractRecently, extraction of blood vessels has aroused widespread interests in medical image analysis. In this work, to accelerate convergence speed and enhance the representation for discriminative features, we introduce the residual block structure in the ResNet into the 3D U-Net, and construct a new 3D Residual U-Net architect to segment the hepatic and portal veins from abdominal CT volumes. In addition, we develop a weighted Dice loss function to cope with the challenges of pixel imbalance, vessel boundary segmentation and small vessels segmentation. Furthermore, based on the prediction results, the post-processing methods of 3D morphological closed operation and volume analysis are employed to smooth the surface of vessels and eliminate noise blocks, respectively. Compared with existing 3D DenseNet, FCN and 3D U-Net, the average Dice coefficients of our method in hepatic veins and portal veins segmentation are 71.7% and 76.5% respectively, which are superior to 55.3% and 53.9% of the 3D DenseNet, 60.2% and 75.6% of the FCN, and 66.4% and 73.9% of the 3D U-Net. Meanwhile, the cross validation results prove that our method is accurate and stable for liver vessel extraction. Bin Fang 0001, Mingqi Gao 0001, Shenhai Zheng, Yi Wang 0074 |
ICIP | 6 |
| 2018 | Image Inpainting via Generative Multi-column Convolutional Neural NetworksabstractIn this paper, we propose a generative multi-column network for image inpainting. This network synthesizes different image components in a parallel manner within one stage. To better characterize global structures, we design a confidence-driven reconstruction loss while an implicit diversified MRF regularization is adopted to enhance local details. The multi-column network combined with the reconstruction and MRF loss propagates local and global information derived from context to the target inpainting regions. Extensive experiments on challenging street view, face, natural objects and scenes manifest that our method produces visual compelling results even without previously common post-processing. Yi Wang 0074, Xin Tao 0001, Xiaojuan Qi 0001, Xiaoyong Shen, Jiaya Jia |
NeurIPS | 1 |
| 2013 | Automatic Multi-Scale Segmentation of Intrahepatic Vessel in CT Images for Liver Surgery PlanningabstractThe processing of blood vessels is an indispensable part in complicated surgeries of livers and hearts as the development of medical image technologies, which requires an automatic segmentation system over CT images of organs. However, the vascular pattern of livers in CT images suffers from low contrast to background so that the existing segmentation technologies are not able to extract the blood vessels completely. In the paper, we propose a new algorithm to extract the blood vessels of livers based on the adaptive multi-scale segmentation. First, we prove that the background histogram of normal scale blood vessels obeys the Gaussian distribution in CT images, and obtain the vascular distribution function from the vascular signal segmented from the background with a local optimal threshold. Second, Hessian matrix is employed to enhance the thin blood vessels before the extraction, and a complete and clear segmentation system for blood vessels is constructed by combining the major and thin blood vessels via filtering. Experimental results show the effectiveness of the proposed method, which is able to extract more complete blood vessels for 3D system, and assist the clinical liver surgeries efficiently. Yi Wang 0074, Bin Fang 0001, Jingrui Pi, Patrick Shen-Pei Wang, Hongguang Wang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | Interconnection of wind farms with grid using a MTDC networkabstractIn the light of the practical project experience, the multi-terminal DC (MTDC) is regarded as one of the preferable solutions to solve the grid interconnection issue of wind generation. This paper mainly focuses on the application of the voltage source converter(VSC) based MTDC technology to integrate large scale wind farms to the electric power grid. A radial MTDC system is explored as the best choice for wind power integration, due to that it can mitigate the fluctuation of the aggregated wind power. Based on the analysis of the VSC model and control, the coordinated control strategy for the proposed MTDC system is designed. The operation performance of a four-terminal MTDC system connecting two DFIG-based wind farms, the local and remote grids is also studied, and the proposed control strategy is adopted to achieve a constant power for long distant transmission to the load center under wind speed variations and faults on DC line. Yi Wang 0074, Yingli Luo, Heming Li, Xiangyu Zhang 0005 |
IECON | 2 |