Dongfang Liu

dblp:51/5217 · DBLP profile ↗
← Back
73ranked-venue papers
9as first author
64since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 8 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 26 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
abstract
Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.
Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Victor Y. Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang
AAAI10
2026 On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning
abstract
Vision-Language-Action models have recently emerged as a powerful paradigm for general-purpose robot learning, enabling agents to map visual observations and natural-language instructions into executable robotic actions. Though popular, they are primarily trained via supervised fine-tuning or training-time reinforcement learning, requiring explicit fine-tuning phases, human interventions, or controlled data collection. Consequently, existing methods remain unsuitable for challenging simulated- or physical-world deployments, where robots must respond autonomously and flexibly to evolving environments. To address this limitation, we introduce a Test-Time Reinforcement Learning for VLAs (TT-VLA), a framework that enables on-the-fly policy adaptation during inference. TT-VLA formulates a dense reward mechanism that leverages step-by-step task-progress signals to refine action policies during test time while preserving the SFT/RL-trained priors, making it an effective supplement to current VLA models. Empirical results show that our approach enhances overall adaptability, stability, and task success in dynamic, previously unseen scenarios under simulated and real-world settings. We believe TT-VLA offers a principled step toward self-improving, deployment-ready VLAs.
Changyu Liu, Yiyang Liu 0003, Taowen Wang, Qiao Zhuang, James Liang, Renjing Xu, Qifan Wang 0001, Dongfang Liu, Cheng Han 0001
ACL (1)9
2026 Ethics of trustworthy AI in healthcare: Challenges, principles, and practical pathways
abstract
Artificial Intelligence (AI) is transforming healthcare by enhancing diagnostics, personalizing treatment planning, and streamlining patient care. Yet, its adoption is hindered by persistent ethical challenges, including algorithmic bias, lack of transparency, privacy risks, and unclear accountability. Existing international frameworks articulate high-level principles but seldom provide operational guidance for clinical deployment. We bridge this gap by synthesizing trust dimensions for healthcare, with measurable metrics for fairness, explainability, privacy, accountability, and robustness, and proposing the Healthcare AI Trustworthiness Index (HAITI), a composite, context-aware readiness score with explicit normalization, weighting, and uncertainty reporting. We outline a development–deployment–governance blueprint and present two case studies (diagnostic bias mitigation; privacy-preserving federated learning). Together, these contributions translate ethical principles into measurable practices that can foster trust, improve equity, and accelerate responsible AI integration in clinical settings.
Pegah Ahadian, Wei Xu 0020, Dongfang Liu, Qiang Guan
Neurocomputing3
2025 MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
abstract
Runjia Zeng, Guangyan Sun, Qifan Wang, Tong Geng, Sohail Dianat, Xiaotian Han, Raghuveer Rao, Xueling Zhang, Cheng Han, Lifu Huang, Dongfang Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Runjia Zeng, Guangyan Sun, Qifan Wang 0001, Tong Geng, Sohail A. Dianat, Raghuveer M. Rao, Xueling Zhang, Cheng Han 0001, Lifu Huang, Dongfang Liu
EMNLP11
2025 Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
abstract
Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic inputs within an end-to-end learning framework. Despite their significant capabilities, VLA models introduce new attack surfaces. This paper systematically evaluates their robustness. Recognizing the unique demands of robotic execution, our attack objectives target the inherent spatial and functional characteristics of robotic systems. In particular, we introduce two untargeted attack objectives that leverage spatial foundations to destabilize robotic actions, and a targeted attack objective that manipulates the robotic trajectory. Additionally, we design an adversarial patch generation approach that places a small, colorful patch within the camera's view, effectively executing the attack in both digital and physical environments. Our evaluation reveals a marked degradation in task success rates, with up to a 100\% reduction across a suite of simulated robotic tasks, highlighting critical security gaps in current VLA architectures. By unveiling these vulnerabilities and proposing actionable evaluation metrics, we advance both the understanding and enhancement of safety for VLA-based robotic systems, underscoring the necessity for continuously developing robust defense strategies prior to physical-world deployments.
Taowen Wang, Cheng Han 0001, James Liang, Dongfang Liu, Luna Xinyu Zhang, Qifan Wang 0001, Jiebo Luo 0001, Ruixiang Tang
ICCV5
2025 Diff-PIC: Revolutionizing Particle-In-Cell Nuclear Fusion Simulation with Diffusion Models
abstract
The rapid development of AI highlights the pressing need for sustainable energy, a critical global challenge for decades. Nuclear fusion, generally seen as a promising solution, has been the focus of intensive research for nearly a century, with investments reaching hundreds of billions of dollars. Recent advancements in Inertial Confinement Fusion (ICF) have drawn significant attention to fusion research, in which Laser-Plasma Interaction (LPI) is critical for ensuring fusion stability and efficiency. However, the complexity of LPI makes analytical approaches impractical, leaving researchers dependent on extremely computationally intensive Particle-in-Cell (PIC) simulations to generate data, posing a significant bottleneck to the advancement of fusion research. In response, this work introduces Diff-PIC, a novel framework that leverages conditional diffusion models as a computationally efficient alternative to PIC simulations for generating high-fidelity scientific LPI data. In this work, physical patterns captured by PIC simulations are distilled into diffusion models associated with two tailored enhancements: (1) To effectively capture the complex relationships between physical parameters and their corresponding outcomes, the parameters are encoded in a physically informed manner. (2) To further enhance efficiency while maintaining physical validity, the rectified flow technique is employed to transform our model into a one-step conditional diffusion model. Experimental results show that Diff-PIC achieves a $\sim$16,200$\times$ speedup compared to traditional PIC on a 100 picosecond simulation, while delivering superior accuracy compared to other data generation approaches.
Chuan Liu 0001, Chunshu Wu, Shihui Cao, Mingkai Chen 0002, James Liang, Ang Li 0006, Chuang Ren, Ying Nian Wu, Dongfang Liu, Tong Geng
ICLR10
2025 Re-Imagining Multimodal Instruction Tuning: A Representation View
abstract
Multimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instruction-following data. However, as the scale of LMMs continues to grow, fully fine-tuning these models has become highly parameter-intensive. Although Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced to reduce the number of tunable parameters, a significant performance gap remains compared to full fine-tuning. Furthermore, existing PEFT approaches are often highly parameterized, making them difficult to interpret and control. In light of this, we introduce Multimodal Representation Tuning (MRT), a novel approach that focuses on directly editing semantically rich multimodal representations to achieve strong performance and provide intuitive control over LMMs. Empirical results show that our method surpasses current state-of-the-art baselines with significant performance gains (e.g., 1580.40 MME score) while requiring substantially fewer tunable parameters (e.g., 0.03% parameters). Additionally, we conduct experiments on editing instrumental tokens within multimodal representations, demonstrating that direct manipulation of these representations enables simple yet effective control over network behavior.
Yiyang Liu 0003, James Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001
ICLR9
2025 DS-LLM: Leveraging Dynamical Systems to Enhance Both Training and Inference of Large Language Models
abstract
The training of large language models (LLMs) faces significant computational cost challenges, limiting their scalability toward artificial general intelligence (AGI) and broader adoption. With model sizes doubling approximately every 3.4 months and training costs escalating from 64 million USD for GPT-4 in 2020 to 191 million USD for Gemini Ultra in 2023, the economic burden has become unsustainable. While techniques such as quantization offer incremental improvements, they fail to address the fundamental computational bottleneck. In this work, we introduce DS-LLM, a novel framework that leverages dynamical system (DS)-based machines, which exploit Natural Annealing to rapidly converge to minimal energy states, yielding substantial efficiency gains. Unlike traditional methods, DS-LLM maps LLM components to optimization problems solvable via Hamiltonian configurations and utilizes continuous electric current flow in DS-machines for hardware-native gradient descent during training. We mathematically demonstrate the equivalence between conventional LLMs and DS-LLMs and present a method for transforming a trained LLM into a DS-LLM. Experimental evaluations across multiple model sizes demonstrate orders-of-magnitude improvements in speed and energy efficiency for both training and inference while maintaining consistent accuracy. Additionally, we provide an in-depth analysis of the challenges and potential solutions associated with this emerging computing paradigm, aiming to lay a solid foundation for future research.
Ruibing Song, Chuan Liu 0001, Chunshu Wu, Ang Li 0006, Dongfang Liu, Ying Nian Wu, Tong Geng
ICLR5
2025 Visual Agents as Fast and Slow Thinkers
abstract
Achieving human-level intelligence requires refining cognitive distinctions between \textit{System 1} and \textit{System 2} thinking. While contemporary AI, driven by large language models, demonstrates human-like traits, it falls short of genuine cognition. Transitioning from structured benchmarks to real-world scenarios presents challenges for visual agents, often leading to inaccurate and overly confident responses. To address the challenge, we introduce \textbf{\textsc{FaST}}, which incorporates the \textbf{Fa}st and \textbf{S}low \textbf{T}hinking mechanism into visual agents. \textsc{FaST} employs a switch adapter to dynamically select between \textit{System 1/2} modes, tailoring the problem-solving approach to different task complexity. It tackles uncertain and unseen objects by adjusting model confidence and integrating new contextual data. With this novel design, we advocate a \textit{flexible system}, \textit{hierarchical reasoning} capabilities, and a \textit{transparent decision-making} pipeline, all of which contribute to its ability to emulate human-like cognitive processes in visual intelligence. Empirical results demonstrate that \textsc{FaST} outperforms various well-known baselines, achieving 80.8\% accuracy over $VQA^{v2}$ for visual question answering and 48.7\% $GIoU$ score over ReasonSeg for reasoning segmentation, demonstrate \textsc{FaST}'s superior performance. Extensive testing validates the efficacy and robustness of \textsc{FaST}'s core components, showcasing its potential to advance the development of cognitive visual agents in AI systems.
Guangyan Sun, Mingyu Jin, Zhenting Wang, Siqi Ma 0005, Qifan Wang 0001, Tong Geng, Ying Nian Wu, Yongfeng Zhang 0003, Dongfang Liu
ICLR10
2025 Robust Ego-Exo Correspondence with Long-Term Memory
abstract
Establishing object-level correspondence between egocentric and exocentric views is essential for intelligent assistants to deliver precise and intuitive visual guidance. However, this task faces numerous challenges, including extreme viewpoint variations, occlusions, and the presence of small objects. Existing approaches usually borrow solutions from video object segmentation models, but still suffer from the aforementioned challenges. Recently, the Segment Anything Model 2 (SAM 2) has shown strong generalization capabilities and excellent performance in video object segmentation. Yet, when simply applied to the ego-exo correspondence (EEC) task, SAM 2 encounters severe difficulties due to ineffective ego-exo feature fusion and limited long-term memory capacity, especially for long videos. Addressing these problems, we propose a novel EEC framework based on SAM 2 with long-term memories by presenting a dual-memory architecture and an adaptive feature routing module inspired by Mixture-of-Experts (MoE). Compared to SAM 2, our approach features **(i)** a Memory-View MoE module which consists of a dual-branch routing mechanism to adaptively assign contribution weights to each expert feature along both channel and spatial dimensions, and **(ii)** a dual-memory bank system with a simple yet effective compression strategy to retain critical long-term information while eliminating redundancy. In the extensive experiments on the challenging EgoExo4D benchmark, our method, dubbed ***LM-EEC***, achieves new state-of-the-art results and significantly outperforms existing methods and the SAM 2 baseline, showcasing its strong generalization across diverse scenarios. Our code and model are available at https://github.com/juneyeeHu/LM-EEC.
Bing Fan, Xin Gu 0003, Haiqing Ren, Dongfang Liu, Heng Fan 0001, Libo Zhang 0001
NeurIPS5
2025 All You Need is One: Capsule Prompt Tuning with a Single Vector
abstract
Prompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning generation with task-aware guidance. Despite its successes, current prompt-based learning methods heavily rely on laborious grid searching for optimal prompt length and typically require considerable number of prompts, introducing additional computational burden. Worse yet, our pioneer findings indicate that the task-aware prompt design is inherently limited by its absence of instance-aware information, leading to a subtle attention interplay with the input sequence. In contrast, simply incorporating instance-aware information as a part of the guidance can enhance the prompt-tuned model performance without additional fine-tuning. Moreover, we find an interesting phenomenon, namely "attention anchor", that incorporating instance-aware tokens at the earliest position of the sequence can successfully preserve strong attention to critical structural information and exhibit more active attention interaction with all input tokens. In light of our observation, we introduce Capsule Prompt-Tuning (CaPT), an efficient and effective solution that leverages off-the-shelf, informative instance semantics into prompt-based learning. Our approach innovatively integrates both instance-aware and task-aware information in a nearly parameter-free manner (i.e., one single capsule prompt). Empirical results demonstrate that our method can exhibit superior performance across various language tasks (e.g., 84.03\% average accuracy on T5-Large), serving as an "attention anchor," while enjoying high parameter efficiency (e.g., 0.003\% of model parameters on Llama3.2-1B).
Yiyang Liu 0003, James Liang, Heng Fan 0001, Yiming Cui 0002, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001
NeurIPS8
2025 Probabilistic Token Alignment for Large Language Model Fusion
abstract
Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more cost-effective alternative is to fuse existing pre-trained LLMs with different architectures into a more powerful model. However, a key challenge in existing model fusion is their dependence on manually predefined vocabulary alignment, which may not generalize well across diverse contexts, leading to performance degradation in several evaluation. To solve this, we draw inspiration from distribution learning and propose the probabilistic token alignment method as a general and soft mapping for alignment, named as PTA-LLM. Our approach innovatively reformulates token alignment into a classic mathematical problem: optimal transport, seamlessly leveraging distribution-aware learning to facilitate more coherent model fusion. Apart from its inherent generality, PTA-LLM exhibits interpretability from a distributional perspective, offering insights into the essence of the token alignment. Empirical results demonstrate that probabilistic token alignment enhances the target model's performance across multiple capabilities.
Runjia Zeng, James Liang, Cheng Han 0001, Zhiwen Cao, Xiaojun Quan, Victor Y. Chen, Lifu Huang, Tong Geng, Qifan Wang 0001, Dongfang Liu
NeurIPS11
2025 Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation
abstract
Semi-supervised learning based on consistency learning offers significant promise for enhancing medical image segmentation. Current approaches use copy-paste as an effective data perturbation technique to facilitate weak-to-strong consistency learning. However, these techniques often lead to a decrease in the accuracy of synthetic labels corresponding to the synthetic data and introduce excessive perturbations to the distribution of the training data. Such over-perturbation causes the data distribution to stray from its true distribution, thereby impairing the model's generalization capabilities as it learns the decision boundaries. We propose a weak-to-strong consistency learning framework that integrally addresses these issues with two primary designs: 1) it emphasizes the use of highly reliable data to enhance the quality of labels in synthetic datasets through cross-copy-pasting between labeled and unlabeled datasets; 2) it employs uncertainty estimation and foreground region constraints to meticulously filter the regions for copy-pasting, thus the copy-paste technique implemented introduces a beneficial perturbation to the training data distribution. Our framework expands the copy-paste method by addressing its inherent limitations, and amplifying the potential of data perturbations for consistency learning. We extensively validated our model using six publicly available medical image segmentation datasets across different diagnostic tasks, including the segmentation of cardiac structures, prostate structures, brain structures, skin lesions, and gastrointestinal polyps. The results demonstrate that our method significantly outperforms state-of-the-art models. For instance, on the PROMISE12 dataset for the prostate structure segmentation task, using only 10% labeled data, our method achieves a 15.31% higher Dice score compared to the baseline models. Our experimental code will be made publicly available at https://github.com/slhuang24/RCP4CL.
Senlong Huang, Yongxin Ge, Dongfang Liu, Mingjian Hong, Junhan Zhao, Alexander C. Loui
IEEE Trans. Image Process.3
2024 KGIF: Optimizing Relation-Aware Recommendations with Knowledge Graph Information Fusion
abstract
While deep-learning-enabled recommender systems demonstrate strong performance benchmarks, many struggle to adapt effectively in real-world environments due to limited use of user-item relationship data and insufficient transparency in recommendation generation. Traditional collaborative filtering approaches fail to integrate multifaceted item attributes, and although Factorization Machines account for item-specific details, they overlook broader relational patterns. Collaborative knowledge graph-based models have progressed by embedding user-item interactions with item-attribute relationships, offering a holistic perspective on interconnected entities. However, these models frequently aggregate attribute and interaction data in an implicit manner, leaving valuable relational nuances underutilized.This study introduces the Knowledge Graph Attention Network with Information Fusion (KGIF), a specialized framework designed to merge entity and relation embeddings explicitly through a tailored self-attention mechanism. The KGIF framework integrates reparameterization via dynamic projection vectors, enabling embeddings to adaptively represent intricate relationships within knowledge graphs. This explicit fusion enhances the interplay between user-item interactions and item-attribute relationships, providing a nuanced balance between user-centric and item-centric representations. An attentive propagation mechanism further optimizes knowledge graph embeddings, capturing multi-layered interaction patterns. The contributions of this work include an innovative method for explicit information fusion, improved robustness for sparse knowledge graphs, and the ability to generate explainable recommendations through interpretable path visualization. The implementation and datasets for this study are publicly available1.
Donghyun Jeon, Houbing Song, Dongfang Liu, Alvaro Velasquez, Chloe Yixin Xie, Shuteng Niu
IEEE Big Data4
2024 ProMotion: Prototypes as Motion Learners
abstract
In this work, we introduce PRoMoTION, a unified proto-typical transformer-based framework engineered to model fundamental motion tasks. PRoMoTION offers a range of compelling attributes that set it apart from current task-specific paradigms. (1) We adopt a prototypical perspective, establishing a unified paradigm that harmonizes disparate motion learning approaches. This novel paradigm stream-lines the architectural design, enabling the simultaneous assimilation of diverse motion information. (2) We capitalize on a dual mechanism involving the feature denoiser and the prototypical learner to decipher the intricacies of motion. This approach effectively circumvents the pitfalls of ambiguity in pixel-wise feature matching, significantly bolstering the robustness of motion representation. (3)) We demon-strate a profound degree of transferability across distinct motion patterns. This inherent versatility reverberates robustly across a comprehensive spectrum of both 2D and 3D downstream tasks. Empirical results demonstrate that PRoMOTION outperforms various well-known specialized architectures, achieving 0.54 and 0.054$AbsRel$error on the Sintel and KITTI depth datasets, 1.04 and 2.01 average endpoint error on the clean and final pass of Sintel flow benchmark, and 4.30 F1-all error on the KITTI flow bench-mark. For its efficacy, we hope our work can catalyze a paradigm shift in universal models in computer vision.
Yawen Lu, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001, Yiming Cui 0002, Zhiwen Cao, Xueling Zhang, Victor Y. Chen, Heng Fan 0001
CVPR2
2024 Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
abstract
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and re-silient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the de-termination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ~ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five bench-mark datasets, including MSRVTT, LSMDC, DiDeMo, VA-TEX, and Charades. Code and models are available here.
Jiamian Wang, Pichao Wang, Dongfang Liu, Sohail A. Dianat, Raghuveer M. Rao, Majid Rabbani, Zhiqiang Tao
CVPR4
2024 Target-driven Attack for Large Language Models
abstract
Current large language models (LLM) provide a strong foundation for large-scale user-oriented natural language tasks. Many users can easily inject adversarial text or instructions through the user interface, thus causing LLM model security challenges like the language model not giving the correct answer. Although there is currently a large amount of research on black-box attacks, most of these black-box attacks use random and heuristic strategies. It is unclear how these strategies relate to the success rate of attacks and thus effectively improve model robustness. To solve this problem, we propose our target-driven black-box attack method to maximize the KL divergence between the conditional probabilities of the clean text and the attack text to redefine the attack’s goal. We transform the distance maximization problem into two convex optimization problems based on the attack goal to solve the attack text and estimate the covariance. Furthermore, the projected gradient descent algorithm solves the vector corresponding to the attack text. Our target-driven black-box attack approach includes two attack strategies: token manipulation and misinformation attack. Experimental results on multiple Large Language Models and datasets demonstrate the effectiveness of our attack method.
Chong Zhang 0006, Mingyu Jin, Dong Shu, Taowen Wang, Dongfang Liu, Xiao-Bo Jin
ECAI5
2024 AMD: Automatic Multi-step Distillation of Large-Scale Vision Models
Cheng Han 0001, Qifan Wang 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Yi Fang 0008, Qiang Guan, Lifu Huang, Dongfang Liu
ECCV (65)9
2024 Radiance Field Learners As UAV First-Person Viewers
Liqi Yan, Qifan Wang 0001, Junhan Zhao, Qiang Guan, Dongfang Liu
ECCV (61)7
2024 M²PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
abstract
Taowen Wang, Yiyang Liu, James Chenhao Liang, Junhan Zhao, Yiming Cui, Yuning Mao, Shaoliang Nie, Jiahao Liu, Fuli Feng, Zenglin Xu, Cheng Han, Lifu Huang, Qifan Wang, Dongfang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Taowen Wang, Yiyang Liu 0003, James Liang, Junhan Zhao, Yiming Cui 0002, Yuning Mao, Shaoliang Nie, Fuli Feng, Zenglin Xu, Cheng Han 0001, Lifu Huang, Qifan Wang 0001, Dongfang Liu
EMNLP14
2024 Fusion Is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection
abstract
Multi-sensor fusion (MSF) is widely used in autonomous vehicles (AVs) for perception, particularly for 3D object detection with camera and LiDAR sensors. The purpose of fusion is to capitalize on the advantages of each modality while minimizing its weaknesses. Advanced deep neural network (DNN)-based fusion techniques have demonstrated the exceptional and industry-leading performance. Due to the redundant information in multiple modalities, MSF is also recognized as a general defence strategy against adversarial attacks. In this paper, we attack fusion models from the camera modality that is considered to be of lesser importance in fusion but is more affordable for attackers. We argue that the weakest link of fusion models depends on their most vulnerable modality and propose an attack framework that targets advanced camera-LiDAR fusion-based 3D object detection models through camera-only adversarial attacks. Our approach employs a two-stage optimization-based strategy that first thoroughly evaluates vulnerable image areas under adversarial attacks, and then applies dedicated attack strategies for different fusion models to generate deployable patches. The evaluations with six advanced camera-LiDAR fusion models and one camera-only model indicate that our attacks successfully compromise all of them. Our approach can either decrease the mean average precision (mAP) of detection performance from 0.824 to 0.353 or degrade the detection score of a target object from 0.728 to 0.156, demonstrating the efficacy of our proposed attack framework. Code is available.
Zhiyuan Cheng 0010, Hongjun Choi, Shiwei Feng 0002, James Liang, Guanhong Tao 0001, Dongfang Liu, Michael Zuzak, Xiangyu Zhang 0001
ICLR6
2024 Image Translation as Diffusion Visual Programmers
abstract
We introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model within the GPT architecture, orchestrating a coherent sequence of visual programs ($i.e.$, computer vision models) for various pro-symbolic steps, which span RoI identification, style transfer, and position manipulation, facilitating transparent and controllable image translation processes. Extensive experiments demonstrate DVP’s remarkable performance, surpassing concurrent arts. This success can be attributed to several key features of DVP: First, DVP achieves condition-flexible translation via instance normalization, enabling the model to eliminate sensitivity caused by the manual guidance and optimally focus on textual descriptions for high-quality content generation. Second, the frame work enhances in-context reasoning by deciphering intricate high-dimensional concepts in feature spaces into more accessible low-dimensional symbols ($e.g.$, [Prompt], [RoI object]), allowing for localized, context-free editing while maintaining overall coherence. Last but not least, DVP improves systemic controllability and explainability by offering explicit symbolic representations at each programming stage, empowering users to intuitively interpret and modify results. Our research marks a substantial step towards harmonizing artificial image translation processes with cognitive intelligence, promising broader applications.
Cheng Han 0001, James Liang, Qifan Wang 0001, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Ying Nian Wu, Dongfang Liu
ICLR8
2024 Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?
abstract
As the scale of vision models continues to grow, the emergence of Visual Prompt Tuning (VPT) as a parameter-efficient transfer learning technique has gained attention due to its superior performance compared to traditional full-finetuning. However, the conditions favoring VPT (the "when") and the underlying rationale (the "why") remain unclear. In this paper, we conduct a comprehensive analysis across 19 distinct datasets and tasks. To understand the "when" aspect, we identify the scenarios where VPT proves favorable by two dimensions: task objectives and data distributions. We find that VPT is preferrable when there is 1) a substantial disparity between the original and the downstream task objectives ($e.g.$, transitioning from classification to counting), or 2) a notable similarity in data distributions between the two tasks ($e.g.$, both involve natural images). In exploring the "why" dimension, our results indicate VPT's success cannot be attributed solely to overfitting and optimization considerations. The unique way VPT preserves original features and adds parameters appears to be a pivotal factor. Our study provides insights into VPT's mechanisms, and offers guidance for its optimal utilization.
Cheng Han 0001, Qifan Wang 0001, Yiming Cui 0002, Wenguan Wang, Lifu Huang, Siyuan Qi, Dongfang Liu
ICLR7
2024 BadPart: Unified Black-box Adversarial Patch Attacks against Pixel-wise Regression Tasks
abstract
Pixel-wise regression tasks (e.g., monocular depth estimation (MDE) and optical flow estimation (OFE)) have been widely involved in our daily life in applications like autonomous driving, augmented reality and video composition. Although certain applications are security-critical or bear societal significance, the adversarial robustness of such models are not sufficiently studied, especially in the black-box scenario. In this work, we introduce the first unified black-box adversarial patch attack framework against pixel-wise regression tasks, aiming to identify the vulnerabilities of these models under query-based black-box attacks. We propose a novel square-based adversarial patch optimization framework and employ probabilistic square sampling and score-based gradient estimation techniques to generate the patch effectively and efficiently, overcoming the scalability problem of previous black-box patch attacks. Our attack prototype, named BadPart, is evaluated on both MDE and OFE tasks, utilizing a total of 7 models. BadPart surpasses 3 baseline methods in terms of both attack performance and efficiency. We also apply BadPart on the Google online service for portrait depth estimation, causing 43.5% relative distance error with 50K queries. State-of-the-art (SOTA) countermeasures cannot defend our attack effectively.
Zhiyuan Cheng 0010, Tengda Guo, Shiwei Feng 0002, Dongfang Liu, MingJie Tang, Xiangyu Zhang 0001
ICML5
2024 Prototypical Transformer As Unified Motion Learners
abstract
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization.
Cheng Han 0001, Yawen Lu, James Liang, Zhiwen Cao, Qifan Wang 0001, Qiang Guan, Sohail A. Dianat, Raghuveer M. Rao, Tong Geng, Zhiqiang Tao, Dongfang Liu
ICML12
2024 SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications
abstract
Communication switches have sometimes been augmented to process collectives, e.g., in the IBM BlueGene and Mellanox SHArP switches. In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third are higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM).
Pouya Haghi, Cheng Tan 0002, Anqi Guo, Chunshu Wu, Dongfang Liu, Ang Li 0006, Anthony Skjellum, Tong Geng, Martin C. Herbordt
ICS5
2024 Efficient Multimodal Semantic Segmentation via Dual-Prompt Learning
abstract
Multimodal (e.g., RGB-Depth/RGB-Thermal) fusion has shown great potential for improving semantic segmentation in complex scenes (e.g., indoor/low-light conditions). Existing approaches often fully fine-tune a dual-branch encoder-decoder framework with a complicated feature fusion strategy for achieving multimodal semantic segmentation, which is training-costly due to the massive parameter updates in feature extraction and fusion. To address this issue, we propose a surprisingly simple yet effective dual-prompt learning network (dubbed DPLNet) for training-efficient multimodal (e.g., RGBD/T) semantic segmentation. The core of DPLNet is to directly adapt a frozen pre-trained RGB model to multimodal semantic segmentation, reducing parameter updates. For this purpose, we present two prompt learning modules, comprising multimodal prompt generator (MPG) and multimodal feature adapter (MFA). MPG works to fuse the features from different modalities in a compact manner and is inserted from shallow to deep stages to generate the multi-level multimodal prompts that are injected into the frozen backbone, while MFA adapts prompted multimodal features in the frozen backbone for better multimodal semantic segmentation. Since both the MPG and MFA are lightweight, only a few trainable parameters (3.88M, 4.4% of the pre-trained backbone parameters) are introduced for multimodal feature fusion and learning. Using a simple decoder (3.27M parameters), DPLNet achieves new state-of-the-art performance or is on a par with other complex approaches on four RGB-D/T semantic segmentation datasets while satisfying parameter efficiency. Moreover, we show DPLNet is general and applicable to other multimodal segmentation tasks. Without special design, DPLNet outperforms many complicated models. The source code can be found at https://github.com/ShaohuaDong2021/DPLNet.
Shaohua Dong, Yunhe Feng, Qing Yang 0003, Yan Huang 0002, Dongfang Liu, Heng Fan 0001
IROS5
2024 Towards Automatic Oracle Prediction for AR Testing: Assessing Virtual Object Placement Quality under Real-World Scenes
abstract
Augmented Reality (AR) technology opens up exciting possibilities in various fields, such as education, work guidance, shopping, communication, and gaming. However, users often encounter usability and user experience issues in current AR apps, often due to the imprecise placement of virtual objects. Detecting these inaccuracies is crucial for AR app testing, but automating the process is challenging due to its reliance on human perception and validation. This paper introduces VOPA (Virtual Object Placement Assessment), a novel approach that automatically identifies imprecise virtual object placements in real-world AR apps. VOPA involves instrumenting real-world AR apps to collect screenshots representing various object placement scenarios and their corresponding metadata under real-world scenes. The collected data are then labeled through crowdsourcing and used to train a hybrid neural network that identifies object placement errors. VOPA aims to enhance AR app testing by automating the assessment of virtual object placement quality and detecting imprecise instances. In our evaluation of a test set of 304 screenshots, VOPA achieved an accuracy of 99.34%, precision of 96.92% and recall of 100%. Furthermore, VOPA successfully identified 38 real-world object placement errors, including instances where objects were hovering between two surfaces or appearing embedded in the wall.
Tahmid Rafi, Dongfang Liu, Xiaoyin Wang, Xueling Zhang
ISSTA4
2024 Tutorial: Large Language-Vision Model in Society
abstract
The tutorial "Large Vision-Language Model in the Society" aims to provide a comprehensive overview of state-of-the-art techniques and applications of large vision-language models (LVLMs), which integrate visual and textual data to transform multimedia research and applications. LVLMs are poised to revolutionize domains such as content creation, social media analysis, education, healthcare, and entertainment by enabling sophisticated content analysis, retrieval, and generation. This tutorial will cover the fundamentals of vision-language integration, state-of-the-art models, training techniques, applications, ethical considerations, and future directions. It is designed to be educational and instructive, providing an in-depth introduction rather than a cursory survey. Attendees will gain practical skills, and insights into the latest research, and engage in interactive sessions to reinforce learning. By addressing both technical and societal aspects, the tutorial will significantly benefit the multimedia community, driving innovation and progress in the field.
Kaicheng Yu, Siyuan Qi, Dongfang Liu
ACM Multimedia4
2024 Diffusion-Inspired Truncated Sampler for Text-Video Retrieval
abstract
Prevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substantial disparities between text and video and can persist in the joint space. In this work, we leverage the potential of Diffusion models to address the text-video modality gap by progressively aligning text and video embeddings in a unified space. However, we identify two key limitations of existing Diffusion models in retrieval tasks: The L2 loss does not fit the ranking problem inherent in text-video retrieval, and the generation quality heavily depends on the varied initial point drawn from the isotropic Gaussian, causing inaccurate retrieval. To this end, we introduce a new Diffusion-Inspired Truncated Sampler (DITS) that jointly performs progressive alignment and modality gap modeling in the joint embedding space. The key innovation of DITS is to leverage the inherent proximity of text and video embeddings, defining a truncated diffusion flow from the fixed text embedding to the video embedding, enhancing controllability compared to adopting the isotropic Gaussian. Moreover, DITS adopts the contrastive loss to jointly consider the relevant and irrelevant pairs, not only facilitating alignment but also yielding a discriminatively structured embedding. Experiments on five benchmark datasets suggest the state-of-the-art performance of DITS. We empirically find that DITS can also improve the structure of the CLIP embedding space. Code is available at https://github.com/Jiamian- Wang/DITS-text-video-retrieval
Jiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
NeurIPS3
2024 Visual Fourier Prompt Tuning
abstract
With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visual prompt tuning is introduced as a parameter-efficient finetuning (PEFT) method to this trend. Despite its successes, a notable research challenge persists within almost all PEFT approaches: significant performance degradation is observed when there is a substantial disparity between the datasets applied in pretraining and finetuning phases. To address this challenge, we draw inspiration from human visual cognition, and propose the Visual Fourier Prompt Tuning (VFPT) method as a general and effective solution for adapting large-scale transformer-based models. Our approach innovatively incorporates the Fast Fourier Transform into prompt embeddings and harmoniously considers both spatial and frequency domain information. Apart from its inherent simplicity and intuitiveness, VFPT exhibits superior performance across all datasets, offering a general solution to dataset challenges, irrespective of data disparities. Empirical results demonstrate that our approach outperforms current state-of-the-art baselines on two benchmarks, with low parameter usage (e.g., 0.57% of model parameters on VTAB-1k) and notable performance enhancements (e.g., 73.20% of mean accuracy on VTAB-1k). Our code is avaliable at https://github.com/runtsang/VFPT.
Runjia Zeng, Cheng Han 0001, Qifan Wang 0001, Chunshu Wu, Tong Geng, Lifu Huang, Ying Nian Wu, Dongfang Liu
NeurIPS8
2024 A systematic evaluation of computational methods for cell segmentation
abstract
Cell segmentation is a fundamental task in analyzing biomedical images. Many computational methods have been developed for cell segmentation and instance segmentation, but their performances are not well understood in various scenarios. We systematically evaluated the performance of 18 segmentation methods to perform cell nuclei and whole cell segmentation using light microscopy and fluorescence staining images. We found that general-purpose methods incorporating the attention mechanism exhibit the best overall performance. We identified various factors influencing segmentation performances, including image channels, choice of training data, and cell morphology, and evaluated the generalizability of methods across image modalities. We also provide guidelines for choosing the optimal segmentation methods in various real application scenarios. We developed Seggal, an online resource for downloading segmentation models already pre-trained with various tissue and cell types, substantially reducing the time and effort for training cell segmentation models.
Junhan Zhao, Hongye Xu, Cheng Han 0001, Zhiqiang Tao, Tong Geng, Dongfang Liu
Briefings Bioinform.8
2024 Self-Supervised Adversarial Training of Monocular Depth Estimation Against Physical-World Attacks
abstract
Monocular Depth Estimation (MDE) plays a vital role in applications such as autonomous driving. However, various attacks target MDE models, with physical attacks posing significant threats to system security. Traditional adversarial training methods, which require ground-truth labels, are not directly applicable to MDE models that lack ground-truth depth. Some self-supervised model hardening techniques (e.g., contrastive learning) overlook the domain knowledge of MDE, resulting in suboptimal performance. In this work, we introduce a novel self-supervised adversarial training approach for MDE models, leveraging view synthesis without the need for ground-truth depth. We enhance adversarial robustness against real-world attacks by incorporating$L_{0}$-norm-bounded perturbation during training. We evaluate our method against supervised learning-based and contrastive learning-based approaches specifically designed for MDE. Our experiments with two representative MDE networks demonstrate improved robustness against various adversarial attacks, with minimal impact on benign performance.
Zhiyuan Cheng 0010, Cheng Han 0001, James Liang, Qifan Wang 0001, Xiangyu Zhang 0001, Dongfang Liu
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Optical Flow as Spatial-Temporal Attention Learners
abstract
Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. To date, the dominant methods are CNN-based, leaving plenty of room for improvement. In this work, we propose TransFlow, a transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and cross-attention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it introduces a concise self-learning paradigm, eliminating the need for complex and laborious multi-stage pre-training procedures. The versatility and superiority of TransFlow extend seamlessly to 3D scene motion, yielding competitive outcomes in 3D scene flow estimation. Our approach attains state-of-the-art results on benchmark datasets such as Sintel and KITTI-15, while also exhibiting exceptional performance on downstream tasks, including video object detection using the ImageNet VID dataset, video frame interpolation using the GoPro dataset, and video stabilization using the DeepStab dataset. We believe that the effectiveness of TransFlow positions it as a flexible baseline for both optical flow and scene flow estimation, offering promising avenues for future research and development.
Yawen Lu, Cheng Han 0001, Qifan Wang 0001, Heng Fan 0001, Zhaodan Kong, Dongfang Liu, Victor Y. Chen
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Guest Editorial Introduction to the Special Issue on Label-Efficient Learning on Video Data
abstract
Currently, the success of image processing relies heavily on large well-annotated datasets. However, collecting and labeling video data are significantly more labor-intensive, posing major challenges for training video algorithms and limiting their practical applications. While label-efficient techniques for image data have advanced, solutions for video data are still emerging. Unlabeled video data, with their inherent structured nature, offer valuable assets for label-efficient learning. Unlike image data, video data naturally captures realistic transformations, providing rich samples for learning. Moreover, from a border perspective, video tasks hold great potential for applications like autonomous driving and video surveillance but present unique challenges due to the need to understand both spatial and temporal aspects. Leveraging label-efficient learning is essential for comprehensively understanding visual content and enabling a wide range of real-world video applications. This Special Issue on “Label-Efficient Learning for Video Data” seeks to advance research in this area, offering new insights and solutions to benefit both researchers and practitioners.
Wenguan Wang, Tianfei Zhou, Dongfang Liu, Zheng Thomas Tang, Alexander C. Loui
IEEE Trans. Circuits Syst. Video Technol.3
2023 MUSTIE: Multimodal Structural Transformer for Web Information Extraction
abstract
Qifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qifan Wang 0001, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu
ACL (1)10
2023 TransFlow: Transformer as Flow Learner
abstract
Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work, we propose TransFlow, a pure transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and crossattention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it enables a concise self-learning paradigm and effectively eliminate the complex and laborious multi-stage pre-training procedures. We achieve the state-of-the-art results on the Sintel, KITTI-15, as well as several downstream tasks, including video object detection, interpolation and stabilization. For its efficacy, we hope TransFlow could serve as a flexible baseline for optical flow estimation.
Yawen Lu, Qifan Wang 0001, Siqi Ma 0005, Tong Geng, Victor Y. Chen, Huaijin G. Chen, Dongfang Liu
CVPR7
2023 APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models
abstract
Qifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Qifan Wang 0001, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu
EMNLP11
2023 Exploiting Logic Locking for a Neural Trojan Attack on Machine Learning Accelerators
abstract
Logic locking has been proposed to safeguard intellectual property (IP) during chip fabrication. Logic locking techniques protect hardware IP by making a subset of combinational modules in a design dependent on a secret key that is withheld from untrusted parties. If an incorrect secret key is used, a set of deterministic errors is produced in locked modules, restricting unauthorized use. A common target for logic locking is neural accelerators, especially as machine-learning-as-a-service becomes more prevalent. In this work, we explore how logic locking can be used to compromise the security of a neural accelerator it protects. Specifically, we show how the deterministic errors caused by incorrect keys can be harnessed to produce neural-trojan-style backdoors. To do so, we first outline a motivational attack scenario where a carefully chosen incorrect key, which we call a trojan key, produces misclassifications for an attacker-specified input class in a locked accelerator. We then develop a theoretically-robust attack methodology to automatically identify trojan keys. To evaluate this attack, we launch it on several locked accelerators. In our largest benchmark accelerator, our attack identified a trojan key that caused a 74% decrease in classification accuracy for attacker-specified trigger inputs, while degrading accuracy by only 1.7% for other inputs on average.
Hongye Xu, Dongfang Liu, Cory E. Merkel, Michael Zuzak
ACM Great Lakes Symposium on VLSI2
2023 E2VPT: An Effective and Efficient Approach for Visual Prompt Tuning
abstract
As the size of transformer-based, models continues to grow, fine-tuning these large-scale pretrained vision models for new tasks has become increasingly parameter-intensive. Parameter-efficient learning has been developed to reduce the number of tunable parameters during fine-tuning. Although these methods show promising results, there is still a significant performance gap compared to full fine-tuning. To address this challenge, we propose an Effective and Efficient Visual Prompt Tuning (E2VPT) approach for large-scale transformer-based model adaptation. Specifically, we introduce a set of learnable key-value prompts and visual prompts into self-attention and input layers, respectively, to improve the effectiveness of model fine-tuning. Moreover, we design a prompt pruning procedure to systematically prune low importance prompts while preserving model performance, which largely enhances the model’s efficiency. Empirical results demonstrate that our approach outperforms several state-of-the-art baselines on two benchmarks, with considerably low parameter usage (e.g., 0.32% of model parameters on VTAB-1k). Our code is available at https://github.com/ChengHan111/E2VPT.
Cheng Han 0001, Qifan Wang 0001, Yiming Cui 0002, Zhiwen Cao, Wenguan Wang, Siyuan Qi, Dongfang Liu
ICCV7
2023 Adversarial Training of Self-supervised Monocular Depth Estimation against Physical-World Attacks
Zhiyuan Cheng 0010, James Liang, Guanhong Tao 0001, Dongfang Liu, Xiangyu Zhang 0001
ICLR4
2023 Visual Recognition with Deep Nearest Centroids
Wenguan Wang, Cheng Han 0001, Tianfei Zhou, Dongfang Liu
ICLR4
2023 CLUSTSEG: Clustering for Universal Segmentation
abstract
We present CLUSTSEG, a general, transformer-based framework that tackles different image segmentation tasks ($i.e.,$ superpixel, semantic, instance, and panoptic) through a unified, neural clustering scheme. Regarding queries as cluster centers, CLUSTSEG is innovative in two aspects: 1) cluster centers are initialized in heterogeneous ways so as to pointedly address task-specific demands ($e.g.,$ instance- or category-level distinctiveness), yet without modifying the architecture; and 2) pixel-cluster assignment, formalized in a cross-attention fashion, is alternated with cluster center update, yet without learning additional parameters. These innovations closely link CLUSTSEG to EM clustering and make it a transparent and powerful framework that yields superior results across the above segmentation tasks.
James Liang, Tianfei Zhou, Dongfang Liu, Wenguan Wang
ICML3
2023 Prompt Learns Prompt: Exploring Knowledge-Aware Generative Prompt Collaboration For Video Captioning
abstract
Fine-tuning large vision-language models is a challenging task. Prompt tuning approaches have been introduced to learn fixed textual or visual prompts while freezing the pre-trained model in downstream tasks. Despite the effectiveness of prompt tuning, what do those learnable prompts learn remains unexplained. In this work, we explore whether prompts in the fine-tuning can learn knowledge-aware prompts from the pre-training, by designing two different sets of prompts in pre-training and fine-tuning phases respectively. Specifically, we present a Video-Language Prompt tuning (VL-Prompt) approach for video captioning, which first efficiently pre-train a video-language model to extract key information (e.g., actions and objects) with flexibly generated Knowledge-Aware Prompt (KAP). Then, we design a Video-Language Prompt (VLP) to transfer the knowledge from the knowledge-aware prompts and fine-tune the model to generate full captions. Experimental results show the superior performance of our approach over several state-of-the-art baselines. We further demonstrate that the video-language prompts are well learned from the knowledge-aware prompts.
Liqi Yan, Cheng Han 0001, Zenglin Xu, Dongfang Liu, Qifan Wang 0001
IJCAI4
2023 ClusterFomer: Clustering As A Universal Visual Learner
abstract
This paper presents ClusterFormer, a universal vision model that is based on the Clustering paradigm with TransFormer. It comprises two novel designs: 1) recurrent cross-attention clustering, which reformulates the cross-attention mechanism in Transformer and enables recursive updates of cluster centers to facilitate strong representation learning; and 2) feature dispatching, which uses the updated cluster centers to redistribute image features through similarity-based metrics, resulting in a transparent pipeline. This elegant design streamlines an explainable and transferable workflow, capable of tackling heterogeneous vision tasks (i.e., image classification, object detection, and image segmentation) with varying levels of clustering granularity (i.e., image-, box-, and pixel-level). Empirical results demonstrate that ClusterFormer outperforms various well-known specialized architectures, achieving 83.41% top-1 acc. over ImageNet-1K for image classification, 54.2% and 47.0% mAP over MS COCO for object detection and instance segmentation, 52.4% mIoU over ADE20K for semantic segmentation, and 55.8% PQ over COCO Panoptic for panoptic segmentation. This work aims to initiate a paradigm shift in universal visual understanding and to benefit the broader field.
James Liang, Yiming Cui 0002, Qifan Wang 0001, Tong Geng, Wenguan Wang, Dongfang Liu
NeurIPS6
2023 Guest Editorial: Learning from limited annotations for computer vision tasks
abstract
The past decade has witnessed remarkable achievements in computer vision, owing to the fast development of deep learning. With the advancement of computing power and deep learning algorithms, we can process and apply millions or even hundreds of millions of large-scale data to train robust and advanced deep learning models. In spite of the impressive success, current deep learning methods tend to rely on massive annotated training data and lack the capability of learning from limited exemplars. However, constructing a million-scale annotated dataset like ImageNet is time-consuming, labour-intensive and even infeasible in many applications. In certain fields, very limited annotated examples can be gathered due to various reasons such as privacy or ethical issues. Consequently, one of the pressing challenges in computer vision is to develop approaches that are capable of learning from limited annotated data. The purpose of this Special Issue is to collect high-quality articles on learning from limited annotations for computer vision tasks (e.g. image classification, object detection, semantic segmentation, instance segmentation and many others), publish new ideas, theories, solutions and insights on this topic and showcase their applications. In this Special Issue we received 29 papers, all of which underwent peer review. Of the 29 originally submitted papers, 9 have been accepted. The nine accepted papers can be clustered into two main categories: theoretical and applications. The papers that fall into the first category are by Liu et al., Li et al. and He et al. The second category of papers offers a direct solution to various computer vision tasks. These papers are by Ma et al., Wu et al., Rao et al., Sun et al., Hou et al. and Gong et al. A brief presentation of each of the papers in this Special Issue follows. Liu et al. present a Gaussianisation prototypical classifier (GPC) for few-shot classification which mainly focuses on solving the issue of prototype bias. GPC consists of handling the features with the Gaussianisation operation and estimating a reliable prototype using the maximum a posteriori method using base class features as prior information. The proposed method is simple yet effective, which does not use any extra labelled data or knowledge. Moreover, it's also a one-step prototype rectification method, which does not resort any complex continuous optimisation. The ablation study shows that GPC can benefit from features pretrained only with CE loss or jointly trained with self-supervised loss. The results demonstrate that the proposed method outperforms related work and other state-of-the-art methods. Li et al. present a novel hyperspectral unmixing method named ‘Global centralised and Structured discriminative Nonnegative Matrix Factorisation (GSNMF)’. The proposed GSNMF offers several distinct advantages over the traditional unmixing techniques. Constructed on the foundation of the manifold regularisation techniques, GSNMF captures the intrinsic structural information by using the local affinity and distant repulsion constraints concurrently. With the structured discriminative information, local affinity constraint ensures that similar elements share similar estimated abundances, while the distant repulsion constraint ensures that dissimilar elements have different abundances. All experiments and analyses have demonstrated that the proposed GSNMF exhibits a remarkable performance compared to the other methods. He et al. present a taxonomy of existing algorithms in the task of makeup transfer. Evaluation methods are proposed, existing methods are analysed and existing datasets are reviewed. Finally the current problems in the field of makeup transfer are discussed, and the trend of future research is analysed. Ma et al. present a dense transformer framework for person re-identification tasks. This paper introduces densely connected class tokens to connect any two layers implicitly. The framework, Denseformer, outperforms other vision transformer models on four widely used benchmarks, namely Market-1501, DukeMTMC-reID, MSMT17 and Occluded-Duke datasets with only a small amount of extra calculation cost. According to the visualisation results, the proposed Denseformer pays more attention to the main parts of human bodies, obtaining discriminative global features. The Denseformer is a general improvement on ViT and works well on other tasks that use ViT as a backbone according to the promising results. Wu et al. present a new homology-continuous-based makeup transformation method, which can be roughly divided into two network branches: the age compensation branch and the makeup transformation branch. Specifically, in the age compensation branch, based on the same source continuity the authors designed a new encoding module which can map the face vector into the corresponding high-dimensional vector space and realise the compensation for age by adjusting the vector direction. In the makeup transformation branch, this work designed a multi-style encoder to handle different types of makeup, such as Chinese Japanese Korean makeup etc. In addition, the proposed network structure is a two-pass encoder-decoder architecture which has good parallelism and can achieve better results with training and inference on GPU. Rao et al. present a novel end-to-end architecture for point completion by using a stack-style folding network called the SSFN. Due to the fact that the output shape code cannot completely represent semantically, they propose a Stack-Style Folding module that transforms the bottleneck output into the style code analogous to StyleGAN. Experiments on ShapeNet and KITTI datasets indicate that the proposed SSFN architecture achieves a decent visual quality and metric performance. Sun et al. present a method for a zero-shot temporal event localisation (ZSTEL) that leverage large-scale video and language models, for example, CLIP. They solve the two key problems for ZSTEL: (1) how to find the relevant region where the event is likely to occur, (2) how to determine event duration after the relevant region is found. They propose the query-guided optimisation for local frame relevance. Relying on the query-to-frame relationship, this method can find the most relevant local frame region where the event is most likely to occur, guided by a constructed objective. The experimental results on the two standard benchmark datasets, Charades-STA and ActivityCaptions have shown the effectiveness of the proposed approach. Hou et al. present a cutting-edge few-shot detection method for logo images. To avoid the misclassification between the base and novel classes, they add an extra classification head. They also apply the convolutional layer into regression heads to improve the accuracy of location by using the limited training data. Considering the characteristics of logo images, they add balanced feature pyramid with Deformable RoI Pooling and unfreeze region proposal network in the fine-tuning stage. The extensive comparative experimentation and ablation studies illustrate the advantage of the proposed method and the effectiveness of every component in the model. Gong et al. present a method for object detection with a long-tail distribution that includes a dual-balanced network and balanced classification loss. This work investigates how the long-tailed distribution impacts the sub-networks in the general two-stage object detection framework Faster-RCNN and finds that unbalanced proposal sampling and unbalanced classification logic deteriorate the performance of the model in terms of AP. They propose the balanced region proposal network and balanced the classification network to address the above issues. Experiments on the LVIS-v0.5 dataset demonstrate that the framework improves the performance of AP without sacrificing too much from the performance of head categories in long-tail distribution. All of the papers selected for this Special Issue show that the field of learning from limited annotations for computer vision tasks is steadily moving forward. The possibility of a weakly supervised learning paradigm will remain a source of inspiration for new techniques in the years to come. Firstly, we wish to express our thanks to Ph.D. students at Nanjing University of Science and Technology for their continuous assistance throughout this process. Also, we wish to express our gratitude to all the contributors who submitted novel scientific results in this special issue and to the anonymous reviewers, whose expert work allowed the realisation of this endeavor. We aspire that this effort should contribute to the further development of DL and increase the concern of the scientific and technological community in the respective area. Last, we should not omit to express our appreciation to the journal's Editors-in-Chief and the Editorial Office for their support throughout this venture. Data sharing is not applicable to this article as no new data were created or analysed in this study. Yazhou Yao is a professor at the School of Computer Science and Engineering and Nanjing University of Science and Technology. With the support of the China Scholarship Council, he received his Ph.D. degree in Computer Science, University of Technology Sydney, Australia at 2018. From July 2018 to July 2019, he worked as a Research Scientist at the Inception Institute of Artificial Intelligence, Abu Dhabi, UAE. His research interests include multimedia processing and machine learning. Wenguan Wang is currently a ZJU100 Young Professor at Zhejiang University. He received his Ph.D. degree from Beijing Institute of Technology in 2018. From 2016 to 2018, he was a joint Ph.D. candidate at the University of California, Los Angeles. From 2018 to 2019, he was a senior scientist at the Inception Institute of Artificial Intelligence, UAE. From 2020 to 2022, he worked as a postdoc researcher at ETH Zurich, Switzerland. After that, he worked as a lecturer and ARC DECRA Fellow at the University of Technology Sydney. His current research interests include computer vision, image processing and deep learning. Qiang Wu received the BEng and MEng degrees in electronic engineering from the Harbin Institute of Technology, Harbin, China, in 1996 and 1998, respectively, and the Ph.D. degree in computing science from the University of Technology Sydney, Sydney, Australia, in 2004. He is currently an Associate Professor and a Core Member of the Global Big Data Technologies Centre, University of Technology Sydney. He has published more than 70 refereed papers, including those published in prestigious journals and top international conferences. His major research interests include computer vision, image processing, pattern recognition, machine learning and multimedia processing. He has served as the chair and/or a Programme Committee Member for a number of international conferences. Dongfang Liu is an Assistant Professor in the Department of Computer Engineering at the Rochester Institute of Technology (RIT). He earned his Ph.D. degree from Purdue University. Dr. Dongfang Liu's research focus on embodied AI and creates general AI solutions to address significant societal challenges. His ongoing work consists of: (1) developing attention-guided perception models that behave like a human's perpetual capacity; and (2) developing structured and human-centred recognition systems that comprehend the surrounding visual world. His publication portfolio includes papers from major conferences in the artificial intelligence and robotics fields, such as CVPR, ECCV, ICCV, ICLR, NIPS, ICML, AAAI, IJCAI, ACL, EMNLP, WWW, WACV, IROS etc. He currently serves on the senior programme committee for AAAI and IJCAI and as an associate editor for IEEE Transactions on Circuits and Systems for Video Technology (TCSVT). Jin Zheng received the BS and MS degrees from Liaoning Technical University, in 2001 and 2004, respectively, and the Ph.D. degree from the School of Computer Science and Engineering, Beihang University, in 2009. She joined the School of Computer Science and Engineering, Beihang University, in 2009. In 2014, she visited Harvard University, MA, USA, as a Visiting Scholar for 1 year. Her current research interests include object detection, tracking and recognition, among other similar interests.
Yazhou Yao, Wenguan Wang, Qiang Wu 0001, Dongfang Liu
IET Comput. Vis.4
2023 Tripartite Feature Enhanced Pyramid Network for Dense Prediction
abstract
Learning pyramidal feature representations is important for many dense prediction tasks (e.g., object detection, semantic segmentation) that demand multi-scale visual understanding. Feature Pyramid Network (FPN) is a well-known architecture for multi-scale feature learning, however, intrinsic weaknesses in feature extraction and fusion impede the production of informative features. This work addresses the weaknesses of FPN through a novel tripartite feature enhanced pyramid network (TFPN), with three distinct and effective designs. First, we develop a feature reference module with lateral connections to adaptively extract bottom-up features with richer details for feature pyramid construction. Second, we design a feature calibration module between adjacent layers that calibrates the upsampled features to be spatially aligned, allowing for feature fusion with accurate correspondences. Third, we introduce a feature feedback module in FPN, which creates a communication channel from the feature pyramid back to the bottom-up backbone and doubles the encoding capacity, enabling the entire architecture to generate incrementally more powerful representations. The TFPN is extensively evaluated over four popular dense prediction tasks, i.e., object detection, instance segmentation, panoptic segmentation, and semantic segmentation. The results demonstrate that TFPN consistently and significantly outperforms the vanilla FPN. Our code is available at https://github.com/jamesliang819.
Dongfang Liu, James Liang, Tong Geng, Alexander C. Loui, Tianfei Zhou
IEEE Trans. Image Process.1
2023 Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence Learning
abstract
Self-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph.
Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui
IEEE Trans. Image Process.3
2022 Towards Unbiased Label Distribution Learning for Facial Pose Estimation Using Anisotropic Spherical Gaussian
Zhiwen Cao, Dongfang Liu, Qifan Wang 0001, Victor Y. Chen
ECCV (12)2
2022 Physical Attack on Monocular Depth Estimation with Optimal Adversarial Patches
Zhiyuan Cheng 0010, James Liang, Hongjun Choi, Guanhong Tao 0001, Zhiwen Cao, Dongfang Liu, Xiangyu Zhang 0001
ECCV (38)6
2022 Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word Generation
abstract
Automatic question generation (AQG) is the task of generating a question from a given passage and an answer.Most existing AQG methods aim at encoding the passage and the answer to generate the question.However, limited work has focused on modeling the correlation between the target answer and the generated question.Moreover, unseen or rare word generation has not been studied in previous works.In this paper, we propose a novel approach which incorporates question generation with its dual problem, question answering, into a unified primal-dual framework.Specifically, the question generation component consists of an encoder that jointly encodes the answer with the passage, and a decoder that produces the question.The question answering component then re-asks the generated question on the passage to ensure that the target answer is obtained.We further introduce a knowledge distillation module to improve the model generalization ability.We conduct an extensive set of experiments on SQuAD and HotpotQA benchmarks.Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods.
Qifan Wang 0001, Xiaojun Quan, Fuli Feng, Dongfang Liu, Zenglin Xu, Sinong Wang, Hao Ma 0001
EMNLP5
2022 GL-RG: Global-Local Representation Granularity for Video Captioning
abstract
Video captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local representation across video frames for caption generation, leaving plenty of room for improvement. In this work, we approach the video captioning task from a new perspective and propose a GL-RG framework for video captioning, namely a Global-Local Representation Granularity. Our GL-RG demonstrates three advantages over the prior efforts: 1) we explicitly exploit extensive visual representations from different video ranges to improve linguistic expression; 2) we devise a novel global-local encoder to produce rich semantic vocabulary to obtain a descriptive granularity of video contents across frames; 3) we develop an incremental training strategy which organizes model learning in an incremental fashion to incur an optimal captioning behavior. Experimental results on the challenging MSR-VTT and MSVD datasets show that our DL-RG outperforms recent state-of-the-art methods by a significant margin. Code is available at https://github.com/ylqi/GL-RG.
Liqi Yan, Qifan Wang 0001, Yiming Cui 0002, Fuli Feng, Xiaojun Quan, Xiangyu Zhang 0001, Dongfang Liu
IJCAI7
2022 Learning Equivariant Segmentation with Instance-Unique Querying
abstract
Prevalent state-of-the-art instance segmentation methods fall into a query-based scheme, in which instance masks are derived by querying the image feature using a set of instance-aware embeddings. In this work, we devise a new training framework that boosts query-based models through discriminative query embedding learning. It explores two essential properties, namely dataset-level uniqueness and transformation equivariance, of the relation between queries and instances. First, our algorithm uses the queries to retrieve the corresponding instances from the whole training dataset, instead of only searching within individual scenes. As querying instances across scenes is more challenging, the segmenters are forced to learn more discriminative queries for effective instance separation. Second, our algorithm encourages both image (instance) representations and queries to be equivariant against geometric transformations, leading to more robust, instance-query matching. On top of four famous, query-based models (i.e., CondInst, SOLOv2, SOTR, and Mask2Former), our training algorithm provides significant performance gains (e.g., +1.6 – 3.2 AP) on COCO dataset. In addition, our algorithm promotes the performance of SOLOv2 by 2.7 AP, on LVISv1 dataset.
Wenguan Wang, James Liang, Dongfang Liu
NeurIPS3
2022 DG-Labeler and DGL-MOTS Dataset: Boost the Autonomous Driving Perception
abstract
Multi-object tracking and segmentation (MOTS) is a critical task for autonomous driving applications. The existing MOTS studies face two critical challenges: 1) the published datasets inadequately capture the real-world complexity for network training to address various driving settings; 2) the working pipeline annotation tool is under-studied in the literature to improve the quality of MOTS learning examples. In this work, we introduce the DG-Labeler and DGL-MOTS dataset to facilitate the training data annotation for the MOTS task and accordingly improve network training accuracy and efficiency. DG-Labeler uses the novel Depth-Granularity Module to depict the instance spatial relations and produce fine-grained instance masks. Annotated by DG-Labeler, our DGL-MOTS dataset exceeds the prior effort (i.e., KITTI MOTS and BDD100K) in data diversity, annotation quality, and temporal representations. Results on extensive cross-dataset evaluations indicate significant performance improvements for several state-of-the-art methods trained on our DGL-MOTS dataset. We believe our DGL-MOTS Dataset and DG-Labeler hold the valuable potential to boost the visual perception of future transportation. Our dataset and code are available here1.
Yiming Cui 0002, Zhiwen Cao, Chloe Yixin Xie, Xingyu Jiang 0001, Feng Tao 0002, Victor Y. Chen, Dongfang Liu
WACV8
2022 WebFormer: The Web-page Transformer for Structure Information Extraction
abstract
Structure information extraction refers to the task of extracting structured text fields from web pages, such as extracting a product offer from a shopping page including product title, description, brand and price. It is an important research topic which has been widely studied in document understanding and web search. Recent natural language models with sequence modeling have demonstrated state-of-the-art performance on web information extraction. However, effectively serializing tokens from unstructured web pages is challenging in practice due to a variety of web layout patterns. Limited work has focused on modeling the web layout for extracting the text fields. In this paper, we introduce WebFormer, a Web-page transFormer model for structure information extraction from web documents. First, we design HTML tokens for each DOM node in the HTML by embedding representations from their neighboring tokens through graph attention. Second, we construct rich attention patterns between HTML tokens and text tokens, which leverages the web layout for effective attention weight computation. We conduct an extensive set of experiments on SWDE and Common Crawl benchmarks. Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods.
Qifan Wang 0001, Yi Fang 0008, Anirudh Ravula, Fuli Feng, Xiaojun Quan, Dongfang Liu
WWW6
2022 In vitro machine learning-based CAR T immunological synapse quality measurements correlate with patient clinical outcomes
abstract
The human immune system consists of a highly intelligent network of billions of independent, self-organized cells that interact with each other. Machine learning (ML) is an artificial intelligence (AI) tool that automatically processes huge amounts of image data. Immunotherapies have revolutionized the treatment of blood cancer. Specifically, one such therapy involves engineering immune cells to express chimeric antigen receptors (CAR), which combine tumor antigen specificity with immune cell activation in a single receptor. To improve their efficacy and expand their applicability to solid tumors, scientists optimize different CARs with different modifications. However, predicting and ranking the efficacy of different "off-the-shelf" immune products (e.g., CAR or Bispecific T-cell Engager [BiTE]) and selection of clinical responders are challenging in clinical practice. Meanwhile, identifying the optimal CAR construct for a researcher to further develop a potential clinical application is limited by the current, time-consuming, costly, and labor-intensive conventional tools used to evaluate efficacy. Particularly, more than 30 years of immunological synapse (IS) research data demonstrate that T cell efficacy is not only controlled by the specificity and avidity of the tumor antigen and T cell interaction, but also it depends on a collective process, involving multiple adhesion and regulatory molecules, as well as tumor microenvironment, spatially and temporally organized at the IS formed by cytotoxic T lymphocytes (CTL) and natural killer (NK) cells. The optimal function of cytotoxic lymphocytes (including CTL and NK) depends on IS quality. Recognizing the inadequacy of conventional tools and the importance of IS in immune cell functions, we investigate a new strategy for assessing CAR-T efficacy by quantifying CAR IS quality using the glass-support planar lipid bilayer system combined with ML-based data analysis. Previous studies in our group show that CAR-T IS quality correlates with antitumor activities in vitro and in vivo. However, current manually quantified IS quality data analysis is time-consuming and labor-intensive with low accuracy, reproducibility, and repeatability. In this study, we develop a novel ML-based method to quantify thousands of CAR cell IS images with enhanced accuracy and speed. Specifically, we used artificial neural networks (ANN) to incorporate object detection into segmentation. The proposed ANN model extracts the most useful information to differentiate different IS datasets. The network output is flexible and produces bounding boxes, instance segmentation, contour outlines (borders), intensities of the borders, and segmentations without borders. Based on requirements, one or a combination of this information is used in statistical analysis. The ML-based automated algorithm quantified CAR-T IS data correlates with the clinical responder and non-responder treated with Kappa-CAR-T cells directly from patients. The results suggest that CAR cell IS quality can be used as a potential composite biomarker and correlates with antitumor activities in patients, which is sufficiently discriminative to further test the CAR IS quality as a clinical biomarker to predict response to CAR immunotherapy in cancer. For translational research, the method developed here can also provide guidelines for designing and optimizing numerous CAR constructs for potential clinical development. Trial Registration: ClinicalTrials.gov NCT00881920.
Alireza Naghizadeh, Wei-chung Tsao, Jong Hyun Cho, Hongye Xu, Mohab Mohamed, Dali Li, Dimitris N. Metaxas, Carlos A. Ramos 0002, Dongfang Liu
PLoS Comput. Biol.10
2022 Video Captioning Using Global-Local Representation
abstract
Video captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local vision representation for sentence generation, leaving plenty of room for improvement. In this work, we approach the video captioning task from a new perspective and propose a GLR framework, namely a global-local representation granularity. Our GLR demonstrates three advantages over the prior efforts. First, we propose a simple solution, which exploits extensive vision representations from different video ranges to improve linguistic expression. Second, we devise a novel global-local encoder, which encodes different video representations including long-range, short-range and local-keyframe, to produce rich semantic vocabulary for obtaining a descriptive granularity of video contents across frames. Finally, we introduce the progressive training strategy which can effectively organize feature learning to incur optimal captioning behavior. Evaluated on the MSR-VTT and MSVD dataset, we outperform recent state-of-the-art methods including a well-tuned SA-LSTM baseline by a significant margin, with shorter training schedules. Because of its simplicity and efficacy, we hope that our GLR could serve as a strong baseline for many video understanding tasks besides video captioning. Code will be available.
Liqi Yan, Siqi Ma 0005, Qifan Wang 0001, Victor Y. Chen, Xiangyu Zhang 0001, Andreas E. Savakis, Dongfang Liu
IEEE Trans. Circuits Syst. Video Technol.7
2021 DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature Aggregation
abstract
In this work, we introduce a Denser Feature Network(DenserNet) for visual localization. Our work provides three principal contributions. First, we develop a convolutional neural network (CNN) architecture which aggregates feature maps at different semantic levels for image representations. Using denser feature maps, our method can produce more key point features and increase image retrieval accuracy. Second, our model is trained end-to-end without pixel-level an-notation other than positive and negative GPS-tagged image pairs. We use a weakly supervised triplet ranking loss to learn discriminative features and encourage keypoint feature repeatability for image representation. Finally, our method is computationally efficient as our architecture has shared features and parameters during forwarding propagation. Our method is flexible and can be crafted on a light-weighted backbone architecture to achieve appealing efficiency with a small penalty on accuracy. Extensive experiment results indicate that our method sets a new state-of-the-art on four challenging large-scale localization benchmarks and three image retrieval benchmarks with the same level of supervision. The code is available at https://github.com/goodproj13/DenserNet
Dongfang Liu, Yiming Cui 0002, Liqi Yan, Christos Mousas, Baijian Yang 0001, Victor Y. Chen
AAAI1
2021 SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
abstract
Video instance segmentation (VIS) is a new and critical task in computer vision. To date, top-performing VIS methods extend the two-stage Mask R-CNN by adding a tracking branch, leaving plenty of room for improvement. In contrast, we approach the VIS task from a new perspective and propose a one-stage spatial granularity network (SG-Net). Compared to the conventional two-stage methods, SG-Net demonstrates four advantages: 1) Our method has a one-stage compact architecture and each task head (detection, segmentation, and tracking) is crafted interdependently so they can effectively share features and enjoy the joint optimization; 2) Our mask prediction is dynamically performed on the sub-regions of each detected instance, leading to high-quality masks of fine granularity; 3) Each of our task predictions avoids using expensive proposal-based RoI features, resulting in much reduced runtime complexity per instance; 4) Our tracking head models objects’ centerness movements for tracking, which effectively enhances the tracking robustness to different object appearances. In evaluation, we present state-of-the-art comparisons on the YouTube-VIS dataset. Extensive experiments demonstrate that our compact one-stage method can achieve improved performance in both accuracy and inference speed. We hope our SG-Net could serve as a strong and flexible base-line for the VIS task. Our code will be available here1.
Dongfang Liu, Yiming Cui 0002, Wenbo Tan, Victor Y. Chen
CVPR1
2021 Hierarchical Attention Fusion for Geo-Localization
abstract
Geo-localization is a critical task in computer vision. In this work, we cast the geo-localization as a 2D image retrieval task. Current state-of-the-art methods for 2D geo-localization are not robust to locate a scene with drastic scale variations because they only exploit features from one semantic level for image representations. To address this limitation, we introduce a hierarchical attention fusion network using multi-scale features for geo-localization. We extract the hierarchical feature maps from a convolutional neural network (CNN) and organically fuse the extracted features for image representations. Our training is self-supervised using adaptive weights to control the attention of feature emphasis from each hierarchical level. Evaluation results on the image retrieval and the large-scale geo-localization benchmarks indicate that our method outperforms the existing state-of-the-art methods. Code is available here: https://github.com/YanLiqi/HAF.
Liqi Yan, Yiming Cui 0002, Victor Y. Chen, Dongfang Liu
ICASSP4
2021 TF-Blender: Temporal Feature Blender for Video Object Detection
abstract
Video objection detection is a challenging task because isolated video frames may encounter appearance deterioration, which introduces great confusion for detection. One of the popular solutions is to exploit the temporal information and enhance per-frame representation through aggregating features from neighboring frames. Despite achieving improvements in detection, existing methods focus on the selection of higher-level video frames for aggregation rather than modeling lower-level temporal relations to increase the feature representation. To address this limitation, we propose a novel solution named TF-Blender, which includes three modules: 1) Temporal relation models the relations between the current frame and its neigh-boring frames to preserve spatial information. 2). Feature adjustment enriches the representation of every neigh-boring feature map; 3) Feature blender combines outputs from the first two modules and produces stronger features for the later detection tasks. For its simplicity, TF-Blender can be effortlessly plugged into any detection network to improve detection behavior. Extensive evaluations on ImageNet VID and YouTube-VIS benchmarks indicate the performance guarantees of using TF-Blender on recent state-of-the-art methods. Code is available at https://github.com/goodproj13/TF-Blender.
Yiming Cui 0002, Liqi Yan, Zhiwen Cao, Dongfang Liu
ICCV4
2021 Semantic Aware Data Augmentation for Cell Nuclei Microscopical Images with Artificial Neural Networks
abstract
There exists many powerful architectures for object detection and semantic segmentation of both biomedical and natural images. However, a difficulty arises in the ability to create training datasets that are large and well-varied. The importance of this subject is nested in the amount of training data that artificial neural networks need to accurately identify and segment objects in images and the infeasibility of acquiring a sufficient dataset within the biomedical field. This paper introduces a new data augmentation method that generates artificial cell nuclei microscopical images along with their correct semantic segmentation labels. Data augmentation provides a step toward accessing higher generalization capabilities of artificial neural networks. An initial set of segmentation objects is used with Greedy AutoAugment to find the strongest performing augmentation policies. The found policies and the initial set of segmentation objects are then used in the creation of the final artificial images. When comparing the state-of-the-art data augmentation methods with the proposed method, the proposed method is shown to consistently outperform current solutions in the generation of nuclei microscopical images.
Alireza Naghizadeh, Hongye Xu, Mohab Mohamed, Dimitris N. Metaxas, Dongfang Liu
ICCV5
2021 A Vector-based Representation to Enhance Head Pose Estimation
abstract
This paper proposes to use the three vectors in a rotation matrix as the representation in head pose estimation and develops a new neural network based on the characteristic of such representation. We address two potential issues existed in current head pose estimation works: 1. Public datasets for head pose estimation use either Euler angles or quaternions to annotate data samples. However, both of these annotations have the issue of discontinuity and thus could result in some performance issues in neural network training. 2. Most research works report Mean Absolute Error (MAE) of Euler angles as the measurement of performance. We show that MAE may not reflect the actual behavior especially for the cases of profile views. To solve these two problems, we propose a new annotation method which uses three vectors to describe head poses and a new measurement Mean Absolute Error of Vectors (MAEV) to assess the performance. We also train a new neural network to predict the three vectors with the constraints of orthogonality. Our proposed method achieves state-of-the-art results on both AFLW2000 and BIWI datasets. Experiments show our vector-based annotation method can effectively reduce prediction errors for large pose angles.
Zhiwen Cao, Zongcheng Chu, Dongfang Liu, Victor Y. Chen
WACV3
2021 Greedy auto-augmentation for n-shot learning using deep neural networks
Alireza Naghizadeh, Dimitris N. Metaxas, Dongfang Liu
Neural Networks3
2020 Object Detection for Autonomous Driving: Motion-aid Feature Calibration Network
Dongfang Liu, Yaqin Mia Wang, Eric T. Matson
ICAART (2)1
2020 Visual Localization for Autonomous Driving: Mapping the Accurate Location in the City Maze
abstract
Accurate localization is a foundational capacity, required for autonomous vehicles to accomplish other tasks such as navigation or path planning. It is a common practice for vehicles to use GPS to acquire location information. However, the application of GPS can result in severe challenges when vehicles run within the inner city where different kinds of structures may shadow the GPS signal and lead to inaccurate location results. To address the localization challenges of urban settings, we propose a novel feature voting technique for visual localization. Different from the conventional front-view-based method, our approach employs views from three directions (front, left, and right) and thus significantly improves the robustness of location prediction. In our work, we craft the proposed feature voting method into three state-of-the-art visual localization networks and modify their architectures properly so that they can be applied for vehicular operation. Extensive field test results indicate that our approach can predict location robustly even in challenging inner-city settings. Our research sheds light on using the visual localization approach to help autonomous vehicles to find accurate location information in a city maze, within a desirable time constraint. The source code is available at github.com/HappyDonkey13/Visual- Localization-for- Autonomous-Driving.
Dongfang Liu, Yiming Cui 0002, Baijian Yang 0001, Victor Y. Chen
ICPR1
2020 A Large-scale Simulation Dataset: Boost the Detection Accuracy for Special Weather Conditions
abstract
Object detection is a fundamental task for autonomous driving systems. One bottleneck hindering detection accuracy is a shortage of well-annotated image data. Virtual reality has provided a feasible low-cost way to facilitate computer vision related developments. In autonomous driving area, existing public datasets from real world generally have data biases and cannot represent a wide range of weather conditions, such as rainy or snowy roads. To address this challenge, we introduce a new large-scale simulation dataset which is generated by an automated pipeline from a high realism video game. Our dataset focuses on weather conditions, which can be adopted to train networks to effectively detect objects under such conditions. We use extensive experiments to evaluate our dataset by comparing it with public datasets. The experiment results show that networks trained with our dataset outperform the networks trained by other public datasets. Our work demonstrates the effectiveness of using simulation data to address real-world challenges in the practice of object detection.
Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen
IJCNN1
2020 Indoor Navigation for Mobile Agents: A Multimodal Vision Fusion Model
abstract
Indoor navigation is a challenging task for mobile agents. The latest vision-based indoor navigation methods make remarkable progress in this field but do not fully leverage visual information for policy learning and struggle to perform well in unseen scenes. To address the existing limitations, we present a multimodal vision fusion model (MVFM). We implement a joint modality of different image recognition networks for navigation policy learning. The proposed model incorporates object detection for target searching, depth estimation for distance prediction, and semantic segmentation to depict the walkable region. In design, our model provides holistic vision knowledge for navigation. Evaluation on AI2-THOR indicates that MVFM improves on the results of a strong baseline model by 3.49% for Success weighted by Path Length (SPL) and 4% for success rate respectively. In comparison with other state-of-the-art systems, MVFM performs in the lead in terms of SPL and success rate. Extensive experiments show the effectiveness of the proposed model.
Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen
IJCNN1
2020 Multimodal Aggregation Approach for Memory Vision-Voice Indoor Navigation with Meta-Learning
abstract
Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal information of visual observation in order to enhance robots' environment understanding. We make use of single RGB images taken by a rst-view monocular camera. We also apply a self-attention mechanism to keep the agent focusing on key areas. Memory is important for the agent to avoid repeating certain tasks unnecessarily and in order for it to adapt adequately to new scenes, therefore, we make use of meta-learning. We have experimented with various functional features extracted from visual observation. Comparative experiments prove that our methods outperform state-of-the-art baselines.
Liqi Yan, Dongfang Liu, Yaoxian Song, Changbin Yu
IROS2
2020 Video object detection for autonomous driving: Motion-aid feature calibration
Dongfang Liu, Yiming Cui 0002, Victor Y. Chen, Jiyong Zhang 0001, Bin Fan 0001
Neurocomputing1
2019 End-to-end Learning Approach for Autonomous Driving: A Convolutional Neural Network Model
Yaqin Mia Wang, Dongfang Liu, Hyewon Jeon, Zhiwei Chu, Eric T. Matson
ICAART (2)2
2019 Virtual World Bridges the Real Challenge: Automated Data Generation for Autonomous Driving
abstract
In autonomous driving research, one of the bottlenecks is the shortage of a well-annotated dataset to train deep neural networks for object detection. Specifically, a dataset focusing on harsh weather conditions is insufficient. The purpose of this research is to explore the power of utilizing synthetic data for training object detection deep neural networks under harsh weather conditions. We introduce a state-of-the-art automated pipeline to collect synthetic images from a high realism video game and generate training data which can be used for training an autonomous driving object detection neural network. We use our synthetic dataset, KITTI, and Cityscapes to train three separate object detection neural networks and employ the PASCAL object detection criteria to evaluate each neural networks' performance. The results from the experiment indicate that the neural network trained by our synthetic dataset outperforms its counterparts and achieves higher average precision (AP) in detecting images under harsh weather conditions. The result sheds a light on employing synthetic data to resolve the challenges in the real world.
Dongfang Liu, Yaqin Mia Wang, Kar Ee Ho, Zhiwei Chu, Eric T. Matson
IV1
2001 OFDM frame synchronization in slotted ALOHA mobile communication systems
abstract
Wireless application of orthogonal frequency division multiplexing (OFDM) is gaining much attention as a modulation technique for high speed data transmission in various applications due to its good properties against frequency selective fading, intersymbol interference (ISI), and intercarrier interference (ICI). However there are still several issues related to synchronization and multiple access that require more research. We investigate the performance of OFDM transmission through slotted ALOHA multiple access and compare the results to unslotted ALOHA OFDM systems. The results show that the OFDM slotted ALOHA technique has advantages in SNR, throughput, and frame synchronization performance compared to unslotted ALOHA with OFDM, which suggests slotted ALOHA as a candidate for multiple access for OFDM wireless communication applications.
Jong-Moon Chung, Dongfang Liu, Krishnaveni Ramasamy, Sriram Varadarajan
VTC Fall2