Dongrui Liu

dblp:199/9200 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 1 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
abstract
Large Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unreliable outputs. However, existing benchmarks fail to reflect such realistic conflict scenarios. Most focus solely on intra-memory conflicts, while context-memory and inter-context conflicts remain largely unaddressed. Furthermore, commonly used factual knowledge-based evaluations are often overlooked, and existing datasets lack a thorough investigation into conflict detection capabilities.To bridge this gap, we propose MMKC-Bench, a benchmark designed to evaluate factual knowledge conflicts in both context-memory and inter-context scenarios. MMKC-Bench encompasses four types of multimodal knowledge conflicts and includes 1,881 knowledge instances and 3,997 images across 32 broad types, collected through automated pipelines with human verification. We evaluate four representative series of LMMs on both model behavior analysis and conflict detection tasks. Our findings show that while current LMMs are capable of recognizing knowledge conflicts, they tend to favor internal parametric knowledge over external evidence. We hope MMKC-Bench will foster further research in multimodal knowledge conflict and enhance the development of multimodal RAG systems.
Yuntao Du 0001, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin 0003, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang 0002, Xiaoye Qu, Qian Li 0043, Dongrui Liu
AAAI14
2026 IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
abstract
Flawed planning from VLM-driven embodied agents poses significant safety hazards, hindering their deployment in real-world household tasks. However, existing static, termination-oriented evaluation paradigms fail to adequately assess risks within these interactive environments, since they cannot simulate dynamic risks that emerge from an agent's actions and rely on unreliable post-hoc evaluations that ignore unsafe intermediate steps. To bridge this critical gap, we propose evaluating an agent's interactive safety: its ability to perceive emergent risks and execute mitigation steps in the correct procedural order. We thus present IS-Bench, the first multi-modal benchmark designed for interactive safety, featuring 161 challenging scenarios with 388 unique safety risks instantiated in a high-fidelity simulator. Crucially, it facilitates a novel process-oriented evaluation that verifies whether risk mitigation actions are performed before/after specific risk-prone steps. Extensive experiments on leading VLMs, including the GPT-4o and Gemini-2.5 series, reveal that current agents lack interactive safety awareness and that while safety-aware Chain-of-Thought can improve performance, it often compromises task completion. By highlighting these critical limitations, IS-Bench provides a foundation for developing safer and more reliable embodied AI systems.
Xiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou, Dongrui Liu, Lu Sheng
AAAI6
2026 AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
abstract
Large Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable capabilities in complex tasks.However, manually designing optimal communication topologies is labor-intensive, while automated expansion methods often result in bloated structures with redundant agents, leading to excessive token consumption.To address this problem, we introduce AgentSlimming, a plug-and-play compression framework for graph-structured multiagent workflows.Motivated by pruning and quantization in neural networks, AgentSlimming compresses workflows by first estimating the importance score of each agent with a hybrid mechanism, and then removes redundant agents or replaces them with low-cost ones, where each operation is validated using a baseline-anchored acceptance rule to prevent performance collapse.Experiments show that AgentSlimming reduces average token cost by up to 78.9% with negligible performance degradation, and sometimes even improves accuracy, achieving a strong Pareto-optimal tradeoff between cost and quality.Our code is publicly available at https://github.com/ CitrusYL/AgentSlimming.
Yulang Chen, Haoxuan Peng, Zichen Wen, Dongrui Liu
ACL (1)5
2026 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
abstract
Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001
ACL (1)20
2026 ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
abstract
Large Reasoning Models (LRMs) with long chain-of-thought reasoning have recently achieved remarkable success.Yet, equipping domain-specialized models with such reasoning capabilities, referred to as "Reasoning + X", remains a significant challenge.While model merging offers a promising training-free solution, existing methods often suffer from a destructive performance collapse: existing methods tend to both weaken reasoning depth and compromise domain-specific utility.Interestingly, we identify a counter-intuitive phenomenon underlying this failure: reasoning ability predominantly resides in parameter regions with low gradient sensitivity, contrary to the common assumption that domain capabilities correspond to high-magnitude parameters.Motivated by this insight, we propose ReasonAny, a novel merging framework that resolves the reasoning-domain performance collapse through Contrastive Gradient Identification.Experiments across safety, biomedicine, and finance domains show that ReasonAny effectively synthesizes "Reasoning + X" capabilities, significantly outperforming state-of-theart baselines while retaining robust reasoning performance.
Junyao Yang, Chen Qian 0010, Yong Liu 0007, Dongrui Liu
ACL (1)6
2026 Entropy-Guided Condensing for Vision Transformer
Sihao Lin, Pumeng Lyu, Dongrui Liu, Zhihui Li 0001, Wenguan Wang, Xiaojun Chang, Yuhui Zheng
Int. J. Comput. Vis.3
2025 VLSBench: Unveiling Visual Leakage in Multimodal Safety
abstract
Safety concerns of Multimodal large language models (MLLMs) have gradually become an important problem in various applications.Surprisingly, previous works indicate a counterintuitive phenomenon that using textual unlearning to align MLLMs achieves comparable safety performances with MLLMs aligned with image-text pairs.To explain such a phenomenon, we discover a Visual Safety Information Leakage (VSIL) problem in existing multimodal safety benchmarks, i.e., the potentially risky content in the image has been revealed in the textual query.Thus, MLLMs can easily refuse these sensitive image-text pairs according to textual queries only, leading to unreliable cross-modality safety evaluation of MLLMs.To this end, we construct multimodal Visual Leakless Safety Bench (VLS-Bench) with 2.2k image-text pairs through an automated data pipeline.Experimental results indicate that VLSBench poses a significant challenge to both open-source and closesource MLLMs, e.g., LLaVA, Qwen2-VL and GPT-4o.Besides, we empirically compare textual and multimodal alignment methods on VLSBench and find that textual alignment is effective enough for multimodal safety scenarios with VSIL, while multimodal alignment is preferable for safety scenarios without VSIL.Code and data are released under https://github.com/AI45Lab/VLSBench.
Xuhao Hu, Dongrui Liu, Xuanjing Huang 0001
ACL (1)2
2025 LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
abstract
Fine-tuning pre-trained Large Language Models (LLMs) for specialized tasks incurs substantial computational and data costs. While model merging offers a training-free solution to integrate multiple task-specific models, existing methods suffer from safety-utility conflicts where enhanced general capabilities degrade safety safeguards. We identify two root causes: $\textbf{neuron misidentification}$ due to simplistic parameter magnitude-based selection, and $\textbf{cross-task neuron interference}$ during merging. To address these challenges, we propose $\textbf{LED-Merging}$, a three-stage framework that $\textbf{L}$ocates task-specific neurons via gradient-based attribution, dynamically $\textbf{E}$lects critical neurons through multi-model importance fusion, and $\textbf{D}$isjoints conflicting updates through parameter isolation. Extensive experiments on Llama-3-8B, Mistral-7B, and Llama2-13B demonstrate that LED-Merging effectively reduces harmful response rates, showing a 31.4\% decrease on Llama-3-8B-Instruct on HarmBench, while simultaneously preserving 95\% of utility performance, such as achieving 52.39\% accuracy on GSM8K. LED-Merging resolves safety-utility conflicts and provides a lightweight, training-free paradigm for constructing reliable multi-task LLMs. Code is available at $\href{https://github.com/MqLeet/LED-Merging}{GitHub}$.
Qianli Ma 0008, Dongrui Liu, Linfeng Zhang 0001
ACL (1)2
2025 The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
abstract
Ensuring awareness of fairness and privacy in Large Language Models (LLMs) is critical.Interestingly, we discover a counterintuitive trade-off phenomenon that enhancing an LLM's privacy awareness through Supervised Fine-Tuning (SFT) methods significantly decreases its fairness awareness with thousands of samples.To address this issue, inspired by the information theory, we introduce a training-free method to Suppress the Privacy and faIrness coupled Neurons (SPIN), which theoretically and empirically decrease the mutual information between fairness and privacy awareness.Extensive experimental results demonstrate that SPIN eliminates the tradeoff phenomenon and significantly improves LLMs' fairness and privacy awareness simultaneously without compromising general capabilities, e.g., improving Qwen-2-7B-Instruct's fairness awareness by 12.2% and privacy awareness by 14.0%.More crucially, SPIN remains robust and effective with limited annotated data or even when only malicious fine-tuning data is available, whereas SFT methods may fail to perform properly in such scenarios.Furthermore, we show that SPIN could generalize to other potential trade-off dimensions.We hope this study provides valuable insights into concurrently addressing fairness and privacy concerns in LLMs and can be integrated into comprehensive frameworks to develop more ethical and responsible AI systems.Our code is available at https://github.com/ChnQ/SPIN.
Chen Qian 0003, Dongrui Liu, Jie Zhang 0121, Yong Liu 0018
ACL (1)2
2025 Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory Perspective
abstract
Despite the remarkable success of attentionbased large language models (LLMs), the precise interaction mechanisms between attention heads remain poorly understood.In contrast to prevalent methods that focus on individual head contributions, we rigorously analyze the intricate interplay among attention heads through a novel framework based on the Harsanyi dividend, a concept from cooperative game theory.Our analysis reveals that significant positive Harsanyi dividends are sparsely distributed across head combinations, indicating that most heads do not contribute cooperatively.Moreover, certain head combinations exhibit negative dividends, indicating implicit competitive relationships.To further optimize the interactions among attention heads, we propose a training-free Game-theoretic Attention Calibration (GAC) method.Specifically, GAC selectively retains heads demonstrating significant cooperative gains and applies fine-grained distributional adjustments to the remaining heads.Comprehensive experiments across 17 benchmarks demonstrate the effectiveness of our proposed GAC and its superior generalization capabilities across diverse model families, scales, and modalities.The source code is available at
Xiaoye Qu, Zengqi Yu, Dongrui Liu, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001
ACL (1)3
2025 LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
abstract
Safety concerns in large language models (LLMs) have gained significant attention due to their exposure to potentially harmful data during pre-training. In this paper, we identify a new safety vulnerability in LLMs: their susceptibility to natural distribution shifts between attack prompts and original toxic prompts, where seemingly benign prompts, semantically related to harmful content, can bypass safety mechanisms. To explore this issue, we introduce a novel attack method, ActorBreaker, which identifies actors related to toxic prompts within pre-training distribution to craft multi-turn prompts that gradually lead LLMs to reveal unsafe content. ActorBreaker is grounded in Latour’s actor-network theory, encompassing both human and non-human actors to capture a broader range of vulnerabilities. Our experimental results demonstrate that ActorBreaker outperforms existing attack methods in terms of diversity, effectiveness, and efficiency across aligned LLMs. To address this vulnerability, we propose expanding safety training to cover a broader semantic space of toxic content. We thus construct a multi-turn safety dataset using ActorBreaker. Fine-tuning models on our dataset shows significant improvements in robustness, though with some trade-offs in utility. Code is available at https://github.com/AI45Lab/ActorAttack.
Qibing Ren, Hao Li 0069, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao 0001, Lei Sha, Junchi Yan, Lizhuang Ma
ACL (1)3
2025 Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
abstract
Diffusion models are trained by learning a sequence of models that reverse each step of noise corruption. Typically, the model parameters are fully shared across multiple timesteps to enhance training efficiency. However, since the denoising tasks differ at each timestep, the gradients computed at different timesteps may conflict, potentially degrading the overall performance of image generation. To solve this issue, this work proposes a Decouple-then-Merge (DeMe) framework, which begins with a pretrained model and finetunes separate models tailored to specific timesteps. We introduce several improved techniques during the fine-tuning stage to promote effective knowledge sharing while minimizing training interference across timesteps. Finally, after finetuning, these separate models can be merged into a single model in the parameter space, ensuring efficient and practical inference. Experimental results show significant generation quality improvements upon 6 benchmarks including Stable Diffusion on COCO30K, ImageNet1K, PartiPrompts, and DDPM on LSUN Church, LSUN Bedroom, and CIFAR10. Code is available at GitHub.
Qianli Ma 0008, Xuefei Ning, Dongrui Liu, Li Niu 0002, Linfeng Zhang 0001
CVPR3
2025 The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
abstract
Estimating the difficulty of input questions as perceived by large language models (LLMs) is essential for accurate performance evaluation and adaptive inference.Existing methods typically rely on repeated response sampling, auxiliary models, or fine-tuning the target model itself, which may incur substantial computational costs or compromise generality.In this paper, we propose a novel approach for difficulty estimation that leverages only the hidden representations produced by the target LLM.We model the token-level generation process as a Markov chain and define a value function to estimate the expected output quality given any hidden state.This allows for efficient and accurate difficulty estimation based solely on the initial hidden state, without generating any output tokens.Extensive experiments across both textual and multimodal tasks demonstrate that our method consistently outperforms existing baselines in difficulty estimation.Moreover, we apply our difficulty estimates to guide adaptive reasoning strategies, including Self-Consistency, Best-of-N, and Self-Refine, achieving higher inference efficiency with fewer generated tokens.
Yubo Zhu, Dongrui Liu, Zecheng Lin, Sheng Zhong 0002
EMNLP2
2025 REEF: Representation Encoding Fingerprints for Large Language Models
abstract
Protecting the intellectual property of open-source Large Language Models (LLMs) is very important, because training LLMs costs extensive computational resources and data. Therefore, model owners and third parties need to identify whether a suspect model is a subsequent development of the victim model. To this end, we propose a training-free REEF to identify the relationship between the suspect and victim models from the perspective of LLMs' feature representations. Specifically, REEF computes and compares the centered kernel alignment similarity between the representations of a suspect model and a victim model on the same samples. This training-free REEF does not impair the model's general capabilities and is robust to sequential fine-tuning, pruning, model merging, and permutations. In this way, REEF provides a simple and effective way for third parties and models' owners to protect LLMs' intellectual property together. Our code is publicly accessible at https://github.com/AI45Lab/REEF.
Jie Zhang 0121, Dongrui Liu, Chen Qian 0010, Linfeng Zhang 0001, Yong Liu 0018, Yu Qiao 0001
ICLR2
2025 Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
abstract
Large reasoning models (LRMs) have demonstrated impressive capabilities in complex problem-solving, yet their internal reasoning mechanisms remain poorly understood. In this paper, we investigate the reasoning trajectories of LRMs from an information-theoretic perspective. By tracking how mutual information (MI) between intermediate representations and the correct answer evolves during LRM reasoning, we observe an interesting MI peaks phenomenon: the MI at specific generative steps exhibits a sudden and significant increase during LRM's reasoning process. We theoretically analyze such phenomenon and show that as MI increases, the probability of model's prediction error decreases. Furthermore, these MI peaks often correspond to tokens expressing reflection or transition, such as "Hmm", "Wait" and "Therefore," which we term as the thinking tokens. We then demonstrate that these thinking tokens are crucial for LRM's reasoning performance, while other tokens has minimal impacts. Building on these analyses, we propose two simple yet effective methods to improve LRM's reasoning performance, by delicately leveraging these thinking tokens. Overall, our work provides novel insights into the reasoning mechanisms of LRMs and offers practical ways to improve their reasoning capabilities. The code is available at \url{https://github.com/ChnQ/MI-Peaks}.
Chen Qian 0010, Dongrui Liu, Haochen Wen, Yong Liu 0018
NeurIPS2
2025 RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
abstract
With the rapid development of multimodal large language models (MLLMs), they are increasingly deployed as autonomous computer-use agents capable of accomplishing complex computer tasks. However, a pressing issue arises: Can the safety risk principles designed and aligned for general MLLMs in dialogue scenarios be effectively transferred to real-world computer-use scenarios? Existing research on evaluating the safety risks of MLLM-based computer-use agents suffers from several limitations: it either lacks realistic interactive environments, or narrowly focuses on one or a few specific risk types. These limitations ignore the complexity, variability, and diversity of real-world environments, thereby restricting comprehensive risk evaluation for computer-use agents. To this end, we introduce **RiOSWorld**, a benchmark designed to evaluate the potential risks of MLLM-based agents during real-world computer manipulations. Our benchmark includes 492 risky tasks spanning various computer applications, involving web, social media, multimedia, os, email, and office software. We categorize these risks into two major classes based on their risk source: (i) User-originated risks and (ii) Environmental risks. For the evaluation, we evaluate safety risks from two perspectives: (i) Risk goal intention and (ii) Risk goal completion. Extensive experiments with multimodal agents on **RiOSWorld** demonstrate that current computer-use agents confront significant safety risks in real-world scenarios. Our findings highlight the necessity and urgency of safety alignment for computer-use agents in real-world computer manipulation, providing valuable insights for developing trustworthy computer-use agents.
Dongrui Liu
NeurIPS3
2024 Explaining Generalization Power of a DNN Using Interactive Concepts
abstract
This paper explains the generalization power of a deep neural network (DNN) from the perspective of interactions. Although there is no universally accepted definition of the concepts encoded by a DNN, the sparsity of interactions in a DNN has been proved, i.e., the output score of a DNN can be well explained by a small number of interactions between input variables. In this way, to some extent, we can consider such interactions as interactive concepts encoded by the DNN. Therefore, in this paper, we derive an analytic explanation of inconsistency of concepts of different complexities. This may shed new lights on using the generalization power of concepts to explain the generalization power of the entire DNN. Besides, we discover that the DNN with stronger generalization power usually learns simple concepts more quickly and encodes fewer complex concepts. We also discover the detouring dynamics of learning complex concepts, which explains both the high learning difficulty and the low generalization power of complex concepts. The code will be released when the paper is accepted.
Huilin Zhou, Hao Zhang 0063, Huiqi Deng, Dongrui Liu, Wen Shen 0002, Shih-Han Chan, Quanshi Zhang
AAAI4
2024 MLP Can Be a Good Transformer Learner
abstract
Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and require same memory costs. This paper introduces a novel strategy that simplifies vision transformers and reduces computational load through the selective removal of non-essential attention layers, guided by entropy considerations. We identify that regarding the attention layer in bottom blocks, their subsequent MLP layers, i.e. two feed-forward layers, can elicit the same entropy quantity. Meanwhile, the accompanied MLPs are under-exploited since they exhibit smaller feature entropy compared to those MLPs in the top blocks. Therefore, we propose to integrate the uninformative attention layers into their subsequent counterparts by degenerating them into identical mapping, yielding only MLP in certain transformer blocks. Experimental results on ImageNet-1k show that the proposed method can remove 40% attention layer of DeiT-B, improving throughput and memory bound without performance compromise.
Sihao Lin, Pumeng Lyu, Dongrui Liu, Xiaodan Liang, Andy Song, Xiaojun Chang
CVPR3
2024 Towards the Dynamics of a DNN Learning Symbolic Interactions
abstract
This study proves the two-phase dynamics of a deep neural network (DNN) learning interactions. Despite the long disappointing view of the faithfulness of post-hoc explanation of a DNN, a series of theorems have been proven [27] in recent years to show that for a given input sample, a small set of interactions between input variables can be considered as primitive inference patterns that faithfully represent a DNN's detailed inference logic on that sample. Particularly, Zhang et al. [41] have observed that various DNNs all learn interactions of different complexities in two distinct phases, and this two-phase dynamics well explains how a DNN changes from under-fitting to over-fitting. Therefore, in this study, we mathematically prove the two-phase dynamics of interactions, providing a theoretical mechanism for how the generalization power of a DNN changes during the training process. Experiments show that our theory well predicts the real dynamics of interactions on different DNNs trained for various tasks.
Qihan Ren, Dongrui Liu, Quanshi Zhang
NeurIPS5
2023 Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different Complexities
abstract
This paper theoretically explains the intuition that simple concepts are more likely to be learned by deep neural networks (DNNs) than complex concepts. In fact, recent studies have observed [24, 15] and proved [26] the emergence of interactive concepts in a DNN, i.e., it is proven that a DNN usually only encodes a small number of interactive concepts, and can be considered to use their interaction effects to compute inference scores. Each interactive concept is encoded by the DNN to represent the collaboration between a set of input variables. Therefore, in this study, we aim to theoretically explain that interactive concepts involving more input variables (i.e., more complex concepts) are more difficult to learn. Our finding clarifies the exact conceptual complexity that boosts the learning difficulty.
Dongrui Liu, Huiqi Deng, Xu Cheng 0005, Qihan Ren, Kangrui Wang, Quanshi Zhang
NeurIPS1
2023 SAKS: Sampling Adaptive Kernels From Subspace for Point Cloud Graph Convolution
abstract
Convolution on 3D point clouds has been extensively explored in geometric deep learning, but it is far from perfect. Convolution operations on point clouds with the fixed kernel indistinguishably capture correspondences between feature pairs, thereby raising an inherent drawback of limited distinctive feature learning. This paper proposes a novel approach to Sampling Adaptive Kernels from Subspace (SAKS) for graph convolution. It adaptively constructs convolution kernels for different feature correspondences according to the unique coordinate representations under the learned subspace. Associating the subspace design with the deep network is a novel concept, providing different viewpoints on feature learning. Specifically, incomplete orthogonal bases are learned at each convolution layer to span a linear subspace in an elaborately designed manner. Subsequently, adaptive kernels are sampled from the learned subspace via unique coordinates parameterized by feature pairs. Unlike existing adaptive convolution methods in a bruteforce manner, the low-rank property of the subspace reduces the computational complexity of this method. Moreover, we theoretically prove that the proposed SAKS derives the principal components of the kernel distribution, which is similar to principal component analysis under some prior assumptions. Extensive experimental results on point cloud classification and segmentation tasks show that SAKS outperforms state-of-the-arts on various benchmark datasets.
Chuanchuan Chen, Dongrui Liu, Trieu-Kien Truong
IEEE Trans. Circuits Syst. Video Technol.2
2023 Self-Supervised Point Cloud Registration With Deep Versatile Descriptors for Intelligent Driving
abstract
As a fundamental yet challenging problem in intelligent transportation systems, point cloud registration attracts vast attention and has been attained with various deep learning-based algorithms. The unsupervised registration algorithms take advantage of deep neural network-enabled novel representation learning while requiring no human annotations, making them applicable to industrial applications. However, unsupervised methods mainly depend on global descriptors, which ignore the high-level representations of local geometries. In this paper, we propose to jointly use both global and local descriptors to register point clouds in a self-supervised manner, which is motivated by a critical observation that all local geometries of point clouds are transformed consistently under the same transformation. Therefore, local geometries can be employed to enhance the representation ability of the feature extraction module. Moreover, the proposed local descriptor is flexible and can be integrated into most existing registration methods and improve their performance. Besides, we also utilize point cloud reconstruction and normal estimation to enhance the transformation awareness of global and local descriptors. Lastly, extensive experimental results on one synthetic and three real-world datasets demonstrate that our method outperforms existing state-of-art unsupervised registration methods and even surpasses supervised ones in some cases. Robustness and computational efficiency evaluations also indicate that the proposed method applies to intelligent vehicles.
Dongrui Liu, Chuanchaun Chen, Robert C. Qiu
IEEE Trans. Intell. Transp. Syst.1
2022 Point Clouds Downsampling Based on Complementary Attention and Contrastive Learning
abstract
This paper presents a novel method for point clouds down-sampling. We formulate the sampling task as an optimal permutation problem and develop two techniques, the com-plementary attention module and contrastive learning mech-anism, to enable it. The complementary attention module assigns large weights to point-wise features to emphasize fea-tures related to selected and discarded points. We optimize the network by introducing the contrastive learning mecha-nism, which minimizes feature discrepancy of the discarded points while maximizing feature separation between the se-lected points. It is evaluated on ModelNet40 and ShapeNet-Core datasets for classification and reconstruction tasks and achieves promising results.
Chuanchuan Chen, Dongrui Liu
IGARSS2
2022 PointFP: A Feature-Preserving Point Cloud Sampling
abstract
Many 3D perception applications (e.g., detection, classification, and segmentation) that directly process point cloud have made great progress with the development of 3D sensors in recent years. However, computational costs and storage demands of these applications grow significantly with the increase of point cloud size. Thus, it is necessary to sam-ple the point cloud while preserving semantic features and uniform density. Deep learning-based sampling methods are task-specific and fail to generalize to different tasks. Hence, we propose a task agnostic point cloud simplification method, called point cloud feature-preserving. The proposed method preserves semantic features from the original point cloud by generating representative nodes. The sampling rate could be controlled by adjusting the number of representative nodes. Qualitative and quantitative experimental results on synthetic and real-world remote sensing point cloud datasets demon-strate the effectiveness of the proposed method.
Dongrui Liu, Chuanchuan Chen, Zhengyun Jiang
IGARSS1
2022 PFMixer: Point Cloud Frequency Mixing
abstract
Convolutional Neural Networks and Transformer are popu-lar for various computer vision tasks. Recently, a modern multi-layer perceptrons (MLPs) architecture, MLP-Mixer, has achieved remarkable performance in visual recognition tasks. MLP-Mixer includes two types of MLPs: channel-mixing and token-mixing MLPs. Inspired by its simplicity and success, we introduce the application of MLP-Mixer for point cloud processing. We design token-mixing operations to allow interaction between high-frequency (edge) regions and low-frequency (smooth) regions. Furthermore, channel-mixing operations are proposed to allow interaction between different channels of local structures and global shapes. Ex-tensive experiments demonstrate the superiority of PFMixer: POINT CLOUD FREQUENCY MIXING
Dongrui Liu, Chuanchuan Chen, Zhengyun Jiang
IGARSS1
2021 Interpreting Representation Quality of DNNs for 3D Point Cloud Processing
abstract
In this paper, we evaluate the quality of knowledge representations encoded in deep neural networks (DNNs) for 3D point cloud processing. We propose a method to disentangle the overall model vulnerability into the sensitivity to the rotation, the translation, the scale, and local 3D structures. Besides, we also propose metrics to evaluate the spatial smoothness of encoding 3D structures, and the representation complexity of the DNN. Based on such analysis, experiments expose representation problems with classic DNNs, and explain the utility of the adversarial training. The code will be released when this paper is accepted.
Wen Shen 0002, Qihan Ren, Dongrui Liu, Quanshi Zhang
NeurIPS3
2021 GeneCGAN: A conditional generative adversarial network based on genetic tree for point cloud reconstruction
Chuanchuan Chen, Dongrui Liu, Trieu-Kien Truong
Neurocomputing2