Wenjing Yang 0002

dblp:48/3396-2 · DBLP profile ↗
← Back
99ranked-venue papers
3as first author
87since 2021 · last 2026
0000-0002-6997-0406ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 3 first-author · 58 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 30 since 2021Systems, architecture and hardware · 9 · 9 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Detecting Unobserved Confounders: A Kernelized Regression Approach
abstract
Detecting unobserved confounders is crucial for reliable causal inference in observational studies. Existing methods require either linearity assumptions or multiple heterogeneous environments, limiting applicability to nonlinear single-environment settings. To bridge this gap, we propose Kernel Regression Confounder Detection (KRCD), a novel method for detecting unobserved confounding in nonlinear observational data under single-environment conditions. KRCD leverages reproducing kernel Hilbert spaces to model complex dependencies. By comparing standard and higher-order kernel regressions, we derive a test statistic whose significant deviation from zero indicates unobserved confounding. Theoretically, we prove two key results: First, in infinite samples, regression coefficients coincide if and only if no unobserved confounders exist. Second, finite-sample differences converge to zero-mean Gaussian distributions with tractable variance. Extensive experiments on synthetic benchmarks and the Twins dataset demonstrate that KRCD not only outperforms existing baselines but also achieves superior computational efficiency.
Yikai Chen, Yunxin Mao, Chunyuan Zheng 0001, Hao Zou 0001, Shanzhi Gu, Yang Shi 0009, Wenjing Yang 0002, Kun Kuang 0001, Haotian Wang 0001
AAAI8
2026 Partial Fairness Awareness: Belief-Guided Strategic Mechanism for Strategic Agents
abstract
Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models. To address fairness concerns intrinsic to strategic classification, recent work has introduced group-specific fairness constraints. However, current fairness-aware approaches face a fundamental dilemma in the issue of fairness exposure: making these constraints public enables strategic manipulation and can lead to fairness reversal, while keeping them hidden may reduce social welfare and discourage genuine improvement. To fill this gap, we subsequently propose the problem of Partial Fairness Awareness (PFA), as our theoretical analysis informs that such a dilemma can be mitigated by releasing the candidate set of fairness constraints and concealing the grounding constraint. To be specific, we introduce a belief-guided strategic mechanism wherein agents iteratively interact with the decision system and maintain a belief distribution over the candidate set of fairness constraints. This belief-guided process enables agents, through iterative interaction and feedback, to update their belief distribution over the candidate set, thereby gradually aligning their belief with the grounding fairness constraint employed by the system. Extensive experiments on real-world and synthetic datasets demonstrate that PFA achieves lower group fairness gaps, higher acceptance of truly qualified individuals, and more stable outcomes compared to fully public or private fairness regimes.
Xinpeng Lv, Chunyuan Zheng 0001, Yunxin Mao, Renzhe Xu, Hao Zou 0001, Shanzhi Gu, Yuanlong Chen, Wenjing Yang 0002, Haotian Wang 0001
AAAI10
2026 ECD: Evidence-guided Contrastive Decoding in Retrieval-Augmented Generation with Accurate Knowledge Reference Adjustment
abstract
Retrieval-Augmented Generation (RAG) enhances the quality of question answering by integrating external knowledge with internal knowledge. A robust RAG system needs to precisely regulate the dependence of the response on the two types of knowledge. The recently proposed context-aware contrastive decoding (CCD) method attempts to achieve this goal by adjusting the knowledge reference weights by comparing the output distribution differences of LLMS when they rely on different knowledge sources. However, these methods are based on probabilistic knowledge reference adjustment strategies (such as the highest probability or entropy), only focus on the relative confidence of the output responses at each decoding step, without considering the absolute confidence of the responses, which may lead to misjudgment of the external knowledge and internal knowledge reference degree in the decoding process. To this end, we propose a novel decoding method, Evidence-guided Contrastive Decoding (ECD), which conducts evidence modeling by constructing the Dirichlet distribution and regards logits as evidence vectors, so as to regulate the reference degree of internal and external knowledge more accurately, and finally improve the quality of generated responses. Extensive evaluations across four public benchmark datasets on three mainstream LLMs have demonstrated the effectiveness and advantages of ECD.
Yize Sui, Wenjing Yang 0002
AAAI5
2026 LCA-Med: A lightweight cross-modal adaptive feature processing module for detecting imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Neural Networks8
2026 A Wolf in Sheep's Clothing: Unveiling a Stealthy Backdoor Attack in Subgraph Federated Learning
abstract
Subgraph Federated Learning (FL) has emerged as a promising paradigm for node classification tasks wherein subgraphs derived from a global graph are distributed across multiple devices to mitigate data leakage risks. Similar to other FL systems, subgraph FL faces significant security challenges, particularly from backdoor attacks, an area that remains extensively underexplored. Existing attacks typically follow a two-phase strategy to implant backdoors. However, in subgraph FL, such attacks often lead toDivergence Amplification, a phenomenon characterized by significant parameter discrepancies between normal and backdoored models, thereby compromising attack stealthiness. To tackle this challenge, we propose BEEF, a Backdoor attack with an End-to-End Framework designed for effectiveness, stealth, and durability. Unlike conventional methods, BEEF incorporates a dedicated trigger generator, which is jointly trained with a backdoored model. To increase its stealthiness, BEEF crafts adversarial perturbations as triggers that provoke misclassification while leaving the model’s parameters entirely untouched. Furthermore, by calibrating a subset of low-salience parameters associated with backdoor activation, BEEF ensures stable performance and sustained effectiveness across FL rounds. Comprehensive evaluations across eight datasets, four models, five state-of-the-art attacks, and six aggregation methods demonstrate BEEF’s effectiveness in deceiving GNNs while maintaining minimal impact on normal data performance. Additionally, we adapt BEEF to federated graph classification tasks, broadening its applicability and practicality.
Hao Yu 0017, Wenjing Yang 0002, Chuan Ma 0001, Lingyuan Meng, Liang Du 0003, Tao Xiang 0001, Xinwang Liu 0002, Kunlun He
IEEE Trans. Inf. Forensics Secur.2
2026 FD-TE Diagnosis: Enhancing Microservice Fault Diagnosis With Frequency Domain Features and Centrality-Aware Time Encoding
abstract
Fault diagnosis in microservice systems requires high availability, driving research towards multimodal learning that leverages heterogeneous monitoring data, including logs, metrics, and traces. These data inherently contain time series and data streams across various modalities. However, existing frameworks fail to exploit temporal dependencies in these dynamic streams effectively. First, existing methods rely too much on original time-domain metrics, which makes it difficult for them to capture periodic patterns and sudden events. Second, most language-bound integrations treat timestamps only as sequential labels or numerical inputs, which ignores the rich contextual information within timestamps. As a result, existing models struggle to comprehend the temporal sequences and data characteristics associated with faults in multimodal data. In this work, we introduce FD-TE Diagnosis, a framework for microservice fault diagnosis that addresses these challenges by applying frequency domain feature analysis and time encoding. To enhance the detection of abnormal events, we integrate a frequency-domain approach that utilizes the Fast Fourier Transform (FFT) to extract robust features from metric data. Next, we encode the timestamps of all events using a dedicated time encoding layer. These temporal representations are then incorporated to strengthen event embedding for fault inference. Experimental results demonstrate that our method enhances the sensitivity of multimodal models in interpreting temporal features, resulting in improved diagnostic outcomes.
Yangfan Li 0001, Haotian Wang 0001, Minglong Li, Fengxiao Tang, Wenjing Yang 0002
IEEE Trans. Reliab.6
2026 AnyUser: Translating Sketched User Intent Into Domestic Robots
abstract
We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.
Songyuan Yang, Huibin Tan, Kailun Yang 0001, Wenjing Yang 0002, Shaowu Yang
IEEE Trans. Robotics4
2025 MRBTP: Efficient Multi-Robot Behavior Tree Planning and Collaboration
abstract
Multi-robot task planning and collaboration are critical challenges in robotics. While Behavior Trees (BTs) have been established as a popular control architecture and are plannable for a single robot, the development of effective multi-robot BT planning algorithms remains challenging due to the complexity of coordinating diverse action spaces. We propose the Multi-Robot Behavior Tree Planning (MRBTP) algorithm, with theoretical guarantees of both soundness and completeness. MRBTP features cross-tree expansion to coordinate heterogeneous actions across different BTs to achieve the team's goal. For homogeneous actions, we retain backup structures among BTs to ensure robustness and prevent redundant execution through intention sharing. While MRBTP is capable of generating BTs for both homogeneous and heterogeneous robot teams, its efficiency can be further improved. We then propose an optional plugin for MRBTP when Large Language Models (LLMs) are available to reason goal-related actions for each robot. These relevant actions can be pre-planned to form long-horizon subtrees, significantly enhancing the planning speed and collaboration efficiency of MRBTP. We evaluate our algorithm in warehouse management and everyday service scenarios. Results demonstrate MRBTP's robustness and execution efficiency under varying settings, as well as the ability of the pre-trained LLM to generate effective task-specific subtrees for MRBTP.
Yishuai Cai, Xinglin Chen, Zhongxuan Cai, Yunxin Mao, Minglong Li, Wenjing Yang 0002, Ji Wang 0001
AAAI6
2025 XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
abstract
The astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the imagery features ultra-high resolution that incorporates extremely complex semantic relationships. Existing benchmarks usually adopt notably smaller image sizes than real-world RS scenarios, suffer from limited annotation quality, and consider insufficient dimensions of evaluation. To address these issues, we present XLRS-Bench: a comprehensive benchmark for evaluating the perception and reasoning capabilities of MLLMs in ultra-high-resolution RS scenarios. XLRS-Bench boasts the largest average image size (8500×8500) observed thus far, with all evaluation samples meticulously annotated manually, assisted by a novel semi-automatic captioner on ultra-high-resolution RS images. On top of the XLRS-Bench, 16 sub-tasks are defined to evaluate MLLMs’ 10 kinds of perceptual capabilities and 6 kinds of reasoning capabilities, with a primary emphasis on advanced cognitive processes that facilitate real-world decision-making and the capture of spatiotemporal changes. The results of both general and RS-focused MLLMs on XLRS-Bench indicate that further efforts are needed for real-world RS applications. We have open-sourced XLRS-Bench to support further research in developing more powerful MLLMs for remote sensing.
Fengxiang Wang 0004, Hongzhen Wang, Zonghao Guo, Di Wang 0023, Yulin Wang 0002, Mingshuo Chen, Long Lan, Wenjing Yang 0002, Jing Zhang 0037, Zhiyuan Liu 0001, Maosong Sun 0001
CVPR9
2025 Anchor-Prompt-based Segmentation and Embedding Model
abstract
Tackling multi-object tracking and segmentation (MOTS) can be attributed to a multi-task learning task, i.e., performing Segmentation and Identity Embedding jointly (SIEJ). Unfortunately, achieving optimal SIEJ is non-trivial, as it relies on different spatiotemporal features of objects. Besides, the lack of labeled data raises the difficulty of balancing the two objectives in SIEJ. Empowered by the Segment Anything Model (SAM), we propose an Anchor-Prompt-based Segmentation and Embedding Model (APSEM) towards optimal SIEJ by introducing redundant anchors and designing an embedding decoder. On the one hand, our APSEM allows redundant anchors for the same object to identify occluded objects and refine the pixel-wise edge mask. On the other hand, our proposed embedding decoder is designed to address identity learning for redundant anchors by minimizing differences between the same identities in redundant prompts. We also create a parameter-efficient fine-tuning strategy, which helps combine the embedding module into the foundation model through a bit of data. Experimental results on MOTSChallenge datasets validate the effectiveness of the proposed APSEM method for MOTS tasks. Such results also demonstrate that each module improves segmentation and identity embedding performance through joint training.
Shuman Li, Haotian Wang 0001, Wenjing Yang 0002, Hengzhu Liu
ICASSP4
2025 Automated Exposure Mapping for Networked Interference
abstract
By characterizing interactions and influences across individuals, networked interference aims to estimate cross-individual treatment effects. For each individual, one of the central components of existing approaches is to manually design an exposure mapping from their neighboring covariates (including their own ones) to different exposure conditions. However, handcraft neighboring structures defined by such manual schemes struggle to capture the complex and flexible structures exhibited by real-world social networks. To bridge this gap, we propose an Automated Exposure Mapping Network (AEMNet) by capturing networked interference conditions automatically with Graph Neural Networks (GNNs) and achieving mapping with deep embedded clustering. The learned representations between individuals in the graph structure reveal patterns and structures hidden behind data, facilitating application on large-scale, relationally complex networked data. We conducted extensive experiments demonstrating that our approach outperforms the baselines in both quality and flexibility, underscoring its ability to better characterize the interference relationships.
Yunxin Mao, Haotian Wang 0001, Yishuai Cai, Minglong Li, Ji Wang 0001, Wenjing Yang 0002
ICASSP6
2025 Robust CLIP-Guided Deep Thinking: A Two-Stage Optimization Strategy for Enhancing Adversarial Robustness and Reliability in LVLMs
abstract
Large Vision-Language models (LVLMs) have demonstrated remarkable performance in a wide range of vision-language tasks as an efficient input/output system. However, the lack of adversarial robustness at the input side and the widespread hallucination phenomenon at the output side significantly undermine user trust in them. Current solutions to the former tend to sacrifice the general performance of LVLMs, while solving the latter requires a large amount of engineering costs. To address these challenges, we propose a two-stage optimization strategy called RCDT (Robust CLIP-guided Deep Thinking), which aims to enhance the adversarial robustness of LVLMs with minimal general performance loss while reducing hallucinations. First, we introduce a constrained adversarial fine-tuning approach for CLIP to limit the general performance loss during the enhancement of robustness. Furthermore, this CLIP is used to think deeply about the output process of LVLMs to reduce hallucinations. Experiments show that RCDT not only reduce general performance loss by more than half while maintaining adversarial robustness compared to the baselines, but also demonstrate good performance in mitigating hallucinations.
Yize Sui, Wanrong Huang, Wenjing Yang 0002, Chaofan Zhao, Ji Wang 0001
ICASSP3
2025 Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
Fengxiang Wang 0004, Hongzhen Wang, Di Wang 0023, Zonghao Guo, Zhenyu Zhong, Long Lan, Wenjing Yang 0002, Jing Zhang 0037
ICCV7
2025 Effective and Efficient Time-Varying Counterfactual Prediction with State-Space Models
abstract
Time-varying counterfactual prediction (TCP) from observational data supports the answer of when and how to assign multiple sequential treatments, yielding importance in various applications. Despite the progress achieved by recent advances, e.g., LSTM or Transformer based causal approaches, their capability of capturing interactions in long sequences remains to be improved in both prediction performance and running efficiency. In parallel with the development of TCP, the success of the state-space models (SSMs) has achieved remarkable progress toward long-sequence modeling with saved running time. Consequently, studying how Mamba simultaneously benefits the effectiveness and efficiency of TCP becomes a compelling research direction. In this paper, we propose to exploit advantages of the SSMs to tackle the TCP task, by introducing a counterfactual Mamba model with Covariate-based Decorrelation towards Selective Parameters (Mamba-CDSP). Motivated by the over-balancing problem in TCP of the direct covariate balancing methods, we propose to de-correlate between the current treatment and the representation of historical covariates, treatments, and outcomes, which can mitigate the confounding bias while preserve more covariate information. In addition, we show that the overall de-correlation in TCP is equivalent to regularizing the selective parameters of Mamba over each time step, which leads our approach to be effective and lightweight. We conducted extensive experiments on both synthetic and real-world datasets, demonstrating that Mamba-CDSP not only outperforms baselines by a large margin, but also exhibits prominent running efficiency.
Haotian Wang 0001, Haoxuan Li 0001, Hao Zou 0001, Haoang Chi, Long Lan, Wanrong Huang, Wenjing Yang 0002
ICLR7
2025 Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation
abstract
The real world naturally has dimensions of time and space. Therefore, estimating the counterfactual outcomes with spatial-temporal attributes is a crucial problem. However, previous methods are based on classical statistical models, which still have limitations in performance and generalization. This paper proposes a novel framework for estimating counterfactual outcomes with spatial-temporal attributes using the Transformer, exhibiting stronger estimation ability. Under mild assumptions, the proposed estimator within this framework is consistent and asymptotically normal. To validate the effectiveness of our approach, we conduct simulation experiments and real data experiments. Simulation experiments show that our estimator has a stronger estimation capability than baseline methods. Real data experiments provide a valuable conclusion to the causal effect of conflicts on forest loss in Colombia. The source code is available at this [URL](https://github.com/lihe-maxsize/DeppSTCI_Release_Version-master).
Haoang Chi, Wanrong Huang, Wenjing Yang 0002
ICML6
2025 HBTP: Heuristic Behavior Tree Planning with Large Language Model Reasoning
abstract
Behavior Trees (BTs) are increasingly becoming a popular control structure in robotics due to their modularity, reactivity, and robustness. In terms of BT generation methods, BT planning shows promise for generating reliable BTs. However, the scalability of BT planning is often constrained by prolonged planning times in complex scenarios, largely due to a lack of domain knowledge. In contrast, pre-trained Large Language Models (LLMs) have demonstrated task reasoning capabilities across various domains, though the correctness and safety of their planning remain uncertain. This paper proposes integrating BT planning with LLM reasoning, introducing Heuristic Behavior Tree Planning (HBTP)-a reliable and efficient framework for BT generation. The key idea in HBTP is to leverage LLMs for task-specific reasoning to generate a heuristic path, which BT planning can then follow to expand efficiently. We first introduce the heuristic BT expansion process, along with two heuristic variants designed for optimal planning and satisficing planning, respectively. Then, we propose methods to address the inaccuracies of LLM reasoning, including action space pruning and reflective feedback, to further enhance both reasoning accuracy and planning efficiency. Experiments demonstrate the theoretical bounds of HBTP, and results from four datasets confirm its practical effectiveness in everyday service robot applications.
Yishuai Cai, Xinglin Chen, Yunxin Mao, Minglong Li, Shaowu Yang, Wenjing Yang 0002, Ji Wang 0001
ICRA6
2025 FutureNet-LoF: Joint Trajectory Prediction and Lane Occupancy Field Prediction with Future Context Encoding
abstract
Most prior motion prediction endeavors in autonomous driving have inadequately encoded future scenarios, leading to predictions that may fail to accurately capture the diverse movements of agents (e.g., vehicles or pedestrians). To address this, we propose FutureNet, which explicitly integrates initially predicted trajectories into the future scenario and further encodes these future contexts to enhance subsequent forecasting. Additionally, most previous motion forecasting works have focused on predicting independent futures for each agent. However, safe and smooth autonomous driving requires accurately predicting the diverse future behaviors of numerous surrounding agents jointly in complex dynamic environments. Given that all agents occupy certain potential travel spaces and possess lane driving priority, we propose Lane Occupancy Field (LOF), a new representation with lane semantics for motion forecasting in autonomous driving. LOF can simultaneously capture the joint probability distribution of all road participants' future spatial-temporal positions. Due to the high compatibility between lane occupancy field prediction and trajectory prediction, we propose a novel network for joint prediction of these two tasks. Our approach ranks 1st on two large-scale motion forecasting benchmarks: Argoverse 1 and Argoverse 2, while it is also the champion method of the CVPR 2024 Argoverse 2 motion forecasting challenge.
Mingkun Wang, Xiaoguang Ren, Ruochun Jin, Minglong Li, Xiaochuan Zhang, Changqian Yu, Wenjing Yang 0002
ICRA8
2025 BTPG: A Platform and Benchmark for Behavior Tree Planning in Everyday Service Robots
abstract
Behavior Trees (BTs) are a widely used control architecture in robotics, renowned for their robustness and safety, which are especially crucial for everyday service robots. Recently, several methods have been proposed to automatically plan BTs to accomplish specific tasks. However, existing research in BT planning lacks two main aspects: (1) the absence of a standard platform for modeling and planning BTs, along with testing benchmarks; and (2) insufficient metrics for a comprehensive evaluation of BT planning algorithms. In this paper, we propose Behavior Tree Planning Gym (BTPG), the first platform and benchmark for BT planning in everyday service robots. In BTPG, behavior nodes are represented by predicate logic, and objects are categorized to better define the predicate domains and action models. The BT planning problem is then formulated in the STRIPS style. We support four environments and three simulators with different action models, which cover most of the needs of everyday service activities. We design a dataset generator for each environment and test three state-of-the-art BT planning algorithms, as well as one proposed by us, using various common metrics. In addition, we design three advanced metrics, planning progress, region distance, and execution robustness, to gain deeper insights into these BT planning algorithms. With a standard test benchmark, we hope BTPG can inspire and accelerate progress in the field of BT planning. Our codes are available at https://github.com/DIDS-EI/BTPG.
Xinglin Chen, Yishuai Cai, Minglong Li, Yunxin Mao, Wenjing Yang 0002, Ji Wang 0001
IJCAI6
2025 UR4NNV: Neural Network Verification, Under-approximation Reachability Works!
abstract
Recently, formal verification of deep neural networks (DNNs) has garnered considerable attention, and over-approximation based methods have become popular due to their effectiveness and efficiency. However, these strategies face challenges in addressing the "unknown dilemma" concerning whether the exact output region or the introduced approximation error violates the property in question. To address this, this paper introduces theUR4NNVverification framework, which utilizes under-approximation reachability analysis for DNN verification for the first timeUR4NNV focuses on DNNs with Rectified Linear Unit (ReLU) activations and employs a binary tree branch-based under-approximation algorithm. In each epoch, UR4NNVunder-approximates a sub-polytope of the reachable set and verifies this polytope against the given property. Through a trial-and-error approach,UR4NNVeffectively falsifies DNN properties while providing confidence levels when reaching verification epoch bounds and failing falsifying properties. Experimental comparisons with existing verification methods demonstrate the effectiveness and efficiency ofUR4NNVsignificantly reducing the impact of the "unknown dilemma".
Taoran Wu, Bai Xue 0001, Ji Wang 0001, Wenjing Yang 0002, Shaojun Deng, Wanwei Liu
IJCNN6
2025 Debiasing Multimodal Large Language Models via Penalization of Language Priors
abstract
In the realms of computer vision and natural language processing, Multimodal Large Language Models (MLLMs) have become indispensable tools, proficient in generating textual responses based on visual inputs. Despite their advancements, our investigation reveals a noteworthy bias: the generated content is often driven more by the inherent priors of the underlying Large Language Models (LLMs) than by the input image. Empirical experiments underscore the persistence of this bias, as MLLMs often provide confident answers even in the absence of relevant images or given incongruent visual inputs. To rectify these biases and redirect the model's focus toward visual information, we propose two simple, training-free strategies. First, for tasks such as classification or multi-choice question answering, we introduce a ''Post-Hoc Debias'' method using an affine calibration step to adjust the output distribution. This approach ensures uniform answer scores when the image is absent, acting as an effective regularization technique to alleviate the influence of LLM priors. For more intricate open-ended generation tasks, we extend this method to ''Visual Debias Decoding'', which mitigates bias by contrasting token log-probabilities conditioned on a correct image versus a meaningless one. Additionally, our investigation sheds light on the instability of MLLMs across various decoding configurations. Through systematic exploration of different settings, we achieve significant performance improvements-surpassing previously reported results-and raise concerns about the fairness of current evaluation practices. Comprehensive experiments substantiate the effectiveness of our proposed strategies in mitigating biases. These strategies not only prove beneficial in minimizing hallucinations but also contribute to the generation of more helpful and precise illustrations.
Yifan Zhang 0004, Yang Shi 0009, Weichen Yu, Qingsong Wen, Xue Wang 0010, Wenjing Yang 0002, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001
ACM Multimedia6
2025 Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
abstract
Long-context video understanding in Multimodal Large Language Models (MLLMs) faces a critical challenge: balancing computational efficiency with the retention of fine-grained spatio-temporal patterns. Existing approaches (e.g., sparse sampling, dense sampling with low resolution, and token compression) suffer from significant information loss in temporal dynamics, spatial details, or subtle interactions, particularly in videos with complex motion or varying resolutions. To address this, we propose Mavors, a novel framework that introduces Multi-granularity video representation for holistic long-video modeling. Specifically, Mavors directly encodes raw video content into latent representations through two core components: 1) an Intra-chunk Vision Encoder (IVE) that preserves high-resolution spatial features via 3D convolutions and Vision Transformers, and 2) an Inter-chunk Feature Aggregator (IFA) that establishes temporal coherence across chunks using transformer-based dependency modeling with chunk-level rotary position encodings. Moreover, the framework unifies image and video understanding by treating images as single-frame videos via sub-image decomposition. Experiments across diverse benchmarks demonstrate Mavors' superiority in maintaining both spatial fidelity and temporal continuity, significantly outperforming existing methods in tasks requiring fine-grained spatio-temporal reasoning.
Yang Shi 0009, Yushuo Guan, Yuanxing Zhang, Weihong Lin, Jingyun Hua, Xinlong Chen, Bohan Zeng, Wentao Zhang 0001, Wenjing Yang 0002, Di Zhang 0026
ACM Multimedia14
2025 Environment Inference for Learning Generalizable Dynamical System
abstract
Data-driven methods offer efficient and robust solutions for analyzing complex dynamical systems but rely on the assumption of I.I.D. data, driving the development of generalization techniques for handling environmental differences. These techniques, however, are limited by their dependence on environment labels, which are often unavailable during training due to data acquisition challenges, privacy concerns, and environmental variability, particularly in large public datasets and privacy-sensitive domains. In response, we propose DynaInfer, a novel method that infers environment specifications by analyzing prediction errors from fixed neural networks within each training round, enabling environment assignments directly from data. We prove our algorithm effectively solves the alternating optimization problem in unlabeled scenarios and validate it through extensive experiments across diverse dynamical systems. Results show that DynaInfer outperforms existing environment assignment techniques, converges rapidly to true labels, and even achieves superior performance when environment labels are available.
Yue He 0001, Haotian Wang 0001, Wenjing Yang 0002, Peng Cui 0001, Zhong Liu 0002
NeurIPS4
2025 Breaking the Gradient Barrier: Unveiling Large Language Models for Strategic Classification
abstract
Strategic classification (SC) explores how individuals or entities modify their features strategically to achieve favorable classification outcomes. However, existing SC methods, which are largely based on linear models or shallow neural networks, face significant limitations in terms of scalability and capacity when applied to real-world datasets with significantly increasing scale, especially in financial services and the internet sector. In this paper, we investigate how to leverage large language models to design a more scalable and efficient SC framework, especially in the case of growing individuals engaged with decision-making processes. Specifically, we introduce GLIM, a gradient-free SC method grounded in in-context learning. During the feed-forward process of self-attention, GLIM implicitly simulates the typical bi-level optimization process of SC, including both the feature manipulation and decision rule optimization. Without fine-tuning the LLMs, our proposed GLIM enjoys the advantage of cost-effective adaptation in dynamic strategic environments. Theoretically, we prove GLIM can support pre-trained LLMs to adapt to a broad range of strategic manipulations. We validate our approach through experiments with a collection of pre-trained LLMs on real-world and synthetic datasets in financial and internet domains, demonstrating that our GLIM exhibits both robustness and efficiency, and offering an effective solution for large-scale SC tasks.
Xinpeng Lv, Yunxin Mao, Haoxuan Li 0001, Ke Liang 0006, Jinxuan Yang, Wanrong Huang, Haoang Chi, Long Lan, Yuanlong Chen, Wenjing Yang 0002, Haotian Wang 0001
NeurIPS11
2025 MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
abstract
Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios.
Yang Shi 0009, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yifan Zhang 0004, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Zhuoran Zhang 0003, Xinlong Chen, Bohan Zeng, Yushuo Guan, Zhang Zhang 0001, Liang Wang 0001, Haoxuan Li 0001, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan 0001, Haotian Wang 0001, Wenjing Yang 0002
NeurIPS23
2025 Elastic Robust Unlearning of Specific Knowledge in Large Language Models
abstract
LLM unlearning aims to remove sensitive or harmful information within the model, thus reducing the potential risk of generating unexpected information. However, existing Preference Optimization (PO)-based unlearning methods suffer two limitations. First, their rigid reward setting limits the effect of unlearning. Second, the lack of robustness causes unlearned information to reappear. To remedy these two weaknesses, we present a novel LLM unlearning optimization framework, namely Elastic Robust Unlearning (ERU), to efficiently and robustly remove specific knowledge from LLMs. We design the elastic reward setting instead of the rigid reward setting to enhance the unlearning performance. Meanwhile, we incorporate the refusal feature ablation into the unlearning process to trigger specific failure patterns for efficiently enhancing the robustness of the PO-based unlearning methods in multiple scenarios. Experimental results show that ERU can improve the unlearning effectiveness significantly while maintaining a high utility performance. Especially, on the WMDP-Bio benchmark, ERU shows a 9\% improvement over the second-best method, and maintains 83\% performance even under 1,000 sample fine-tuned retraining attacks, significantly better than the baseline method.
Yize Sui, Wenjing Yang 0002, Ruochun Jin, Xiyao Liu
NeurIPS3
2025 GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
abstract
Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce **SuperRS-VQA** (avg. 8,376$\times$8,376) and **HighRS-VQA** (avg. 2,000$\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: *Background Token Pruning* and *Anchored Token Selection*, to reduce the memory footprint while preserving key semantics. Integrating these techniques, we introduce **GeoLLaVA-8K**, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench. Datasets and code were released at https://github.com/MiliLab/GeoLLaVA-8K.
Fengxiang Wang 0004, Mingshuo Chen, Di Wang 0023, Haotian Wang 0001, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang 0002, Hongzhen Wang, Wenjing Yang 0002, Bo Du 0001, Jing Zhang 0037
NeurIPS12
2025 RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
abstract
Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. While the linear-complexity Mamba architecture offers a promising alternative, existing RS applications of Mamba remain limited to supervised tasks on small, domain-specific datasets. To address these challenges, we propose RoMA, a framework that enables scalable self-supervised pretraining of Mamba-based RS foundation models using large-scale, diverse, unlabeled data. RoMA enhances scalability for high-resolution images through a tailored auto-regressive learning strategy, incorporating two key innovations: 1) a rotation-aware pretraining mechanism combining adaptive cropping with angular embeddings to handle sparsely distributed objects with arbitrary orientations, and 2) multi-scale token prediction objectives that address the extreme variations in object scales inherent to RS imagery. Systematic empirical studies validate that Mamba adheres to RS data and parameter scaling laws, with performance scaling reliably as model and data size increase. Furthermore, experiments across scene classification, object detection, and semantic segmentation tasks demonstrate that RoMA-pretrained Mamba models consistently outperform ViT-based counterparts in both accuracy and computational efficiency. The source code and pretrained models have be released at https://github.com/MiliLab/RoMA.
Fengxiang Wang 0004, Yulin Wang 0002, Mingshuo Chen, Haotian Wang 0001, Hongzhen Wang, Haiyan Zhao 0001, Yangang Sun, Di Wang 0023, Long Lan, Wenjing Yang 0002, Jing Zhang 0037
NeurIPS11
2025 Advanced Strategic Improvement with Decision Interactions
Wenjing Yang 0002, Xinpeng Lv, Yunxin Mao, Ruochun Jin, Jinxuan Yang, Yuanlong Chen, Haotian Wang 0001
ECML/PKDD (1)1
2025 Enhancing Uncertainty Quantification in Large Language Models through Semantic Graph Density
abstract
Large Language Models (LLMs) excel in language understanding but are susceptible to "confabulation," where they generate arbitrary, factually incorrect responses to uncertain questions. Detecting confabulation in question answering often relies on Uncertainty Quantification (UQ), which measures semantic entropy or consistency among sampled answers. While several methods have been proposed for UQ in LLMs, they suffer from key limitations, such as overlooking fine-grained semantic relationships among answers and neglecting answer probabilities. To address these issues, we propose Semantic Graph Density (SGD). SGD quantifies semantic consistency by evaluating the density of a semantic graph that captures fine-grained semantic relationships among answers. Additionally, it integrates answer probabilities to adjust the contribution of each edge to the overall uncertainty score. We theoretically prove that SGD generalizes the previous state-of-the-art method, Deg, and empirically demonstrate its superior performance across four LLMs and four free-form question-answering datasets. In particular, in experiments with Llama3.1-8B, SGD outperformed the best baseline by 1.52% in AUROC on the CoQA dataset and by 1.22% in AUARC on the TriviaQA dataset.
Zhaoye Li, Wenjing Yang 0002, Ruochun Jin, Ligong Cao
UAI3
2025 Learning Feasible Causal Algorithmic Recourse: A Prior Structural Knowledge Free Approach
abstract
Algorithmic recourse (AR) has made significant progress by identifying small perturbations in input features that can alter predictions, which provide a data-centric approach to understand decisions from diverse black-box models on the Web. Towards the feasibility issue, i.e., whether the recoursed examples provides actionable and reliable recommendations to end-users, causal algorithmic recourse have incorporated structural causal model (SCM) to preserve the realistic constraints among input features. For instance, preserving structural causal knowledge between "age" and "educational level" can avoid generating samples with decreasing age and increasing educational level. However, previous causal AR methods suffer from the requirement of prior structural causal knowledge, e.g., prior causal graph or the whole SCM, which restricts the realistic application of causal AR methods.
Haotian Wang 0001, Hao Zou 0001, Xueguang Zhou, Shangwen Wang, Wenjing Yang 0002, Peng Cui 0001
WWW5
2025 NT-FAN: A simple yet effective noise-tolerant few-shot adaptation network
Wenjing Yang 0002, Haoang Chi, Yibing Zhan, Xiaoguang Ren, Dapeng Tao, Long Lan
Artif. Intell.1
2025 A visual state space Model-Based Cross-Domain adaptive detection method for imbalanced medical image distribution
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Hudan Pan, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Appl. Intell.8
2025 Dragon Boat Optimization: A Meta-Heuristic for Intelligent Systems
abstract
ABSTRACT Dragon boat racing, a popular aquatic folklore team sport, is traditionally held during the Dragon Boat Festival. Inspired by this event, we propose a novel human‐based meta‐heuristic algorithm called dragon boat optimization (DBO) in this paper. It models the unique behaviours of each crew member on the dragon boat during the race by introducing social psychology mechanisms (social loafing, social incentive). Throughout this process, the focus is on the interaction and collaboration among the crew members, as well as their decision‐making in various situations. During each iteration, DBO implements different state updating strategies. By accurately modelling the crew's behaviour and employing adaptive state update strategies, DBO consistently achieves high optimization performance, as validated by comprehensive testing on 29 benchmark functions and 2 structural design problems. Experimental results indicate that DBO outperforms 7 and 16 state‐of‐the‐art meta‐heuristic algorithms across these test functions and problems, respectively.
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Expert Syst. J. Knowl. Eng.6
2025 Self-supervised re-identification for online joint multi-object tracking
abstract
Recently, the bottleneck of multi-object tracking is shifting from detection performance to association performance. However, research on association algorithms requires a large number of identity labels, which are more expensive than detection labels. To circumvent the need for identity labels, we propose a Self-supervised Re-identification module for online joint Multi-Object Tracking (SR-MOT). Specifically, we design an appearance discriminator to judge identities based solely on detection hypotheses and then associate the same identity with the final trajectory. To train the discriminator without using identity labels, we construct negative pairs by the detections that appear in the same video frame, as they definitely belong to different identities. Positive pairs are naturally constructed through several useful data augmentation strategies at the box level. In addition, our proposed method balances conflicting detection and re-ID tasks by using different output features and dynamically adjusts detection and re-ID loss weights based on the information content of the loss distribution to promote balance between the two tasks from the feature level and optimization methods. In our evaluation on the MOT Challenge benchmark, we show that our SR-MOT performs comparably to supervised methods and is significantly superior to other unsupervised methods. Our proposed method provides a practical solution for multi-object tracking without the need for identity labels, making it more accessible for real-world applications.
Shuman Li, Longqi Yang 0002, Huibin Tan, Binglin Wang, Wanrong Huang, Hengzhu Liu, Wenjing Yang 0002, Long Lan
Knowl. Inf. Syst.7
2025 BIRDNN: Behavior-Imitation Based Repair for Deep Neural Networks
Taoran Wu, Changyuan Zhao, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002, Ji Wang 0001, Wanrong Huang
Neural Networks6
2025 WildVideo: Benchmarking LMMs for Understanding Video-Language Interaction
abstract
We introduce WildVideo, an open-world benchmark dataset designed to address how to assess hallucination of Large Multi-modal Models (LMMs) for understanding video-language interaction in the wild. Our WildVideo comprehensively tests the perceptual, cognitive, and contextual comprehension hallucination of LMMs through both single-turn and multi-turn open-ended question-answering (QA) tasks on videos captured from two human perspectives (i.e. first-person view and third-person view). We define 9 distinct tasks that challenge LMMs across multi-level perceptual tasks (e.g., static and dynamic perception), multi-aspect cognitive tasks (e.g., commonsense, world knowledge), and multi-faceted contextual comprehension tasks (e.g., contextual ellipsis, cross-turn retrieval). The benchmark consists of 1,318 meticulously curated videos, supplemented with 13,704 single-turn QA pairs and 1,585 multi-turn dialogues (up to 5 turns). We evaluated 14 commonly-used LMMs on WildVideo, revealing significant hallucination issues of current LMMs, highlighting substantial gaps in their current capabilities.
Songyuan Yang, Weijiang Yu, Wenjing Yang 0002, Xinwang Liu 0002, Huibin Tan, Long Lan, Nong Xiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Eyes on Islanded Nodes: Better Reasoning via Structure Augmentation and Feature Co-Training on Bi-Level Knowledge Graphs
abstract
Knowledge graphs (KGs) represent known entities and their relationships using triplets, but this method cannot represent relationships between facts, limiting their expressiveness. Recently, the Bi-level Knowledge Graph (Bi-level KG) has addressed this issue by modeling facts as nodes and establishing relationships between these facts, introducing two new tasks: triplet prediction and conditional link prediction. Existing methods enhance triplets through data augmentation method and represent facts using entity representations. However, these methods do not address the isolated nodes at the structure level, nor do they effectively capture the information of facts at the feature level. To address these two issues, we design a data augmentation method that identifies islanded node by detecting anomalous structures and features in the graph. Subsequently, we perform similar subgraph matching for each isolated node to construct potential facts. To enrich the features of facts, we design a weighted combination initialization method for facts and introduce a new relation $\widetilde {R}$ , to connect facts with related entities. This approach allows for the co-training of fact and entity representations during the training process. Extensive experiments validate the effectiveness of our data augmentation and co-training methods. Our model achieves optimal performance in triplet prediction and conditional link prediction tasks.
Hao Li 0146, Ke Liang 0006, Wenjing Yang 0002, Lingyuan Meng, Sihang Zhou 0001, Xinwang Liu 0002
IEEE Trans. Image Process.3
2024 Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer Tracking
abstract
Regarded as a template-matching task for a long time, visual object tracking has witnessed significant progress in space-wise exploration. However, since tracking is performed on videos with substantial time-wise information, it is important to simultaneously mine the temporal contexts which have not yet been deeply explored. Previous supervised works mostly consider template reform as the breakthrough point, but they are often limited by additional computational burdens or the quality of chosen templates. To address this issue, we propose a Space-Time Consistent Transformer Tracker (STCFormer), which uses a sequential fusion framework with multi-granularity consistency constraints to learn spatiotemporal context information. We design a sequential fusion framework that recombines template and search images based on tracking results from chronological frames, fusing updated tracking states in training. To further overcome the over-reliance on the fixed template without increasing computational complexity, we design three space-time consistent constraints: Label Consistency Loss (LCL) for label-level consistency, Attention Consistency Loss (ACL) for patch-level ROI consistency, and Semantic Consistency Loss (SCL) for feature-level semantic consistency. Specifically, in ACL and SCL, the label information is used to constrain the attention and feature consistency of the target and the background, respectively, to avoid mutual interference. Extensive experiments have shown that our STCFormer outperforms many of the best-performing trackers on several popular benchmarks.
Wenjing Yang 0002, Wanrong Huang, Xianchen Zhou, Mingyu Cao, Huibin Tan
AAAI2
2024 Scaling Few-Shot Learning for the Open World
abstract
Few-shot learning (FSL) aims to enable learning models with the ability to automatically adapt to novel (unseen) domains in open-world scenarios. Nonetheless, there exists a significant disparity between the vast number of new concepts encountered in the open world and the restricted available scale of existing FSL works, which primarily focus on a limited number of novel classes. Such a gap hinders the practical applicability of FSL in realistic scenarios. To bridge this gap, we propose a new problem named Few-Shot Learning with Many Novel Classes (FSL-MNC) by substantially enlarging the number of novel classes, exceeding the count in the traditional FSL setup by over 500-fold. This new problem exhibits two major challenges, including the increased computation overhead during meta-training and the degraded classification performance by the large number of classes during meta-testing. To overcome these challenges, we propose a Simple Hierarchy Pipeline (SHA-Pipeline). Due to the inefficiency of traditional protocols of EML, we re-design a lightweight training strategy to reduce the overhead brought by much more novel classes. To capture discriminative semantics across numerous novel classes, we effectively reconstruct and leverage the class hierarchy information during meta-testing. Experiments show that the proposed SHA-Pipeline significantly outperforms not only the ProtoNet baseline but also the state-of-the-art alternatives across different numbers of novel classes.
Wenjing Yang 0002, Haotian Wang 0001, Haoang Chi, Long Lan, Ji Wang 0001
AAAI2
2024 Fake Node-Based Perception Poisoning Attacks against Federated Object Detection Learning in Mobile Computing Networks
abstract
Federated learning (FL) supports massive edge devices to collaboratively train object detection models in mobile computing scenarios. However, the distributed nature of FL exposes significant security vulnerabilities. Existing attack methods either require considerable costs to compromise the majority of participants, or suffer from poor attack success rates. Inspired by this, we devise an efficient fake node-based perception poisoning attacks strategy (FNPPA) to target such weaknesses. In particular, FNPPA poisons local data and injects multiple fake nodes to participate in aggregation, aiming to make the local poisoning model more likely to overwrite clean updates. Moreover, it can achieve greater malicious influence on target objects at a lower cost without affecting the normal detection of other objects. We demonstrate through exhaustive experiments that FNPPA exhibits superior attack impact than the state-of-the-art in terms of average precision and aggregation effect.
Mingxing Duan, Zhuo Tang, Wenjing Yang 0002
DAC5
2024 Diversifying Cross-Domain Few-Shot Learning via Multimodal Image Editing
abstract
Standing out as one of the most widely used tools in Cross-Domain Few-Shot Learning (CDFSL), data augmentation forms the bedrock of numerous recent advancements. However, the current augmentations in CDFSL are limited in their ability to modify high-level semantic attributes, resulting in a lack of diversity along key semantic dimensions. One of the most promising tools to edit images with key semantic attributes, e.g. backgrounds, is image-to-image generation via large multimodal models (LMMs). Given the promising image editing results of recent LMMs, we delve into leveraging LMMs to augment data diversity for CDFSL. We propose a novel method named, Multimodal Few-shot Image Editing (MFIE), which uses LMMs to automatically translate class-specific images into class-agnostic natural language descriptions for various key semantic attributes in target domains and editing origin images based on class-agnostic natural language descriptions. To filter out corrupted data that disturbs the class-specific information, we apply semantic filtering using image-language similarity. Experiments on Meta-Datset show that MFIE surpasses SOTA CDFSL algorithms.
Wenjing Yang 0002, Long Lan, Mingyang Geng, Haotian Wang 0001, Haoang Chi, Xueqiong Li, Ji Wang 0001
ICASSP2
2024 CRNet: Cross-Reconstruction Network for Inconsistent Point Cloud Registration
abstract
Deep learning methods have made significant advancements in point cloud registration, achieving excellent performance on consistent point clouds. However, these methods face challenges when dealing with point clouds exhibiting inconsistent spatial distributions. To address this issue, we propose the Cross-Reconstruction Network (CRNet), a novel approach designed to register two inconsistent point clouds by reconstructing them into a consistent shape. CRNet utilizes a cross-learning framework that facilitates feature interaction between input point clouds at both global and point-wise levels. This interaction network enables the bidirectional generation of corresponding points to reconstruct consistent point clouds for transformation estimation. Furthermore, the transformation parameters can be refined by a regression network to achieve more accurate registration. The experimental results valuated on benchmark datasets demonstrate that CRNet outperforms state-of-the-art methods in inconsistent scenarios.
Yunzhe Xiao, Xueqiong Li, Shaowu Yang, Wenjing Yang 0002, Yong Dou
ICME4
2024 Task Allocation in Heterogeneous Multi-Robot Systems Based on Preference-Driven Hedonic Game
abstract
Multiple preferences between robots and tasks have been largely overlooked in previous research on Multi-Robot Task Allocation (MRTA) problems. In this paper, we propose a preference-driven approach based on hedonic game to address the task allocation problem of muti-robot systems in emergency rescue scenarios. We present a distributed framework considering various preferences between robots and tasks to determine the division of coalitions in such problems and evaluate the scalability and adaptability of our algorithm through relevant experiments. Furthermore, considering the strict communication limitations in emergency rescue scenarios, we have verified that our algorithm can efficiently converge to a Nash-stable coalition partition even in conditions of insufficient communication distance.
Liwang Zhang, Minglong Li, Wenjing Yang 0002, Shaowu Yang
ICRA3
2024 Integrating Intent Understanding and Optimal Behavior Planning for Behavior Tree Generation from Human Instructions
Xinglin Chen, Yishuai Cai, Yunxin Mao, Minglong Li, Wenjing Yang 0002, Ji Wang 0001
IJCAI5
2024 Coalition Formation Game Approach for Task Allocation in Heterogeneous Multi-Robot Systems under Resource Constraints
abstract
This paper studies a case of the multi-robot task allocation (MRTA) problem, where each unmanned aerial vehicle (UAV) is endowed with multiple but limited resources. Completing each task necessitates UAVs to combine different resources through coalition formation, which will incur various costs including flight cost, execution cost, and cooperation cost. To minimize the total cost while maximizing both task completion rate and resource utilization rate, we model the MRTA problem of the UAVs as a leader-follower coalition formation game. In this game, leader UAVs coordinate follower UAVs to fulfill task resource requisites. Meanwhile, follower UAVs select suitable coalitions to join based on the altruistic preference. Theoretical analysis confirms the existence of a Nash stable partition in the coalition formation game. To achieve this stable partition, we propose a coalition formation algorithm. Simulation experiments validate that the proposed algorithm outperforms existing methods for the MRTA problem under resource constraints in terms of both task completion rate and resource utilization rate.
Liwang Zhang, Minglong Li, Wenjing Yang 0002, Shaowu Yang
IROS4
2024 Your Neighbor Matters: Towards Fair Decisions Under Networked Interference
abstract
In the era of big data, decision-making in social networks may introduce bias due to interconnected individuals. For instance, in peer-to-peer loan platforms on the Web, considering an individual's attributes along with those of their interconnected neighbors, including sensitive attributes, is vital for loan approval or rejection downstream. Unfortunately, conventional fairness approaches often assume independent individuals, overlooking the impact of one person's sensitive attribute on others' decisions. To fill this gap, we introduce "Interference-aware Fairness" (IAF) by defining two forms of discrimination as Self-Fairness (SF) and Peer-Fairness (PF), leveraging advances in interference analysis within causal inference. Specifically, SF and PF causally capture and distinguish discrimination stemming from an individual's sensitive attributes (with fixed neighbors' sensitive attributes) and from neighbors' sensitive attributes (with fixed self's sensitive attributes), separately. Hence, a network-informed decision model is fair only when SF and PF are satisfied simultaneously, as interventions in individuals' sensitive attributes or those of their peers both yield equivalent outcomes. To achieve IAF, we develop a deep doubly robust framework to estimate and regularize SF and PF metrics for decision models. Extensive experiments on synthetic and real-world datasets validate our proposed concepts and methods.
Wenjing Yang 0002, Haotian Wang 0001, Haoxuan Li 0001, Hao Zou 0001, Ruochun Jin, Kun Kuang 0001, Peng Cui 0001
KDD1
2024 Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
abstract
Causal reasoning capability is critical in advancing large language models (LLMs) towards artificial general intelligence (AGI). While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality and providing responses that obey the laws of causality, it remains unclear whether they perform genuine causal reasoning akin to humans. However, current evidence indicates the contrary. Specifically, LLMs are only capable of performing shallow (level-1) causal reasoning, primarily attributed to the causal knowledge embedded in their parameters, but they lack the capacity for genuine human-like (level-2) causal reasoning. To support this hypothesis, methodologically, we delve into the autoregression mechanism of transformer-based LLMs, revealing that it is not inherently causal. Empirically, we introduce a new causal Q&A benchmark named CausalProbe 2024, whose corpus is fresh and nearly unseen for the studied LLMs. Empirical results show a significant performance drop on CausalProbe 2024 compared to earlier benchmarks, indicating that LLMs primarily engage in level-1 causal reasoning.To bridge the gap towards level-2 causal reasoning, we draw inspiration from the fact that human reasoning is usually facilitated by general knowledge and intended goals. Inspired by this, we propose G$^2$-Reasoner, a LLM causal reasoning method that incorporates general knowledge and goal-oriented prompts into LLMs' causal reasoning processes. Experiments demonstrate that G$^2$-Reasoner significantly enhances LLMs' causal reasoning capability, particularly in fresh and fictitious contexts. This work sheds light on a new path for LLMs to advance towards genuine causal reasoning, going beyond level-1 and making strides towards level-2.
Haoang Chi, Wenjing Yang 0002, Feng Liu 0003, Long Lan, Xiaoguang Ren, Tongliang Liu, Bo Han 0003
NeurIPS3
2024 DBTN: An adaptive neural network for multiple-disease detection via imbalanced medical images distribution
Xiang Li 0089, Long Lan, Chang-Yong Sun, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Heng Liu 0001, Yudong Zhang 0001
Appl. Intell.6
2024 EAFP-Med: An efficient adaptive feature processing module based on prompts for medical image detection
abstract
The rapid proliferation of medical imaging technologies presents a significant challenge for cross-domain adaptive image detection, as lesion representations can vary dramatically across technologies. To address this issue, we draw inspiration from large language models to propose EAFP-Med, an efficient adaptive feature processing module based on prompts for medical image detection. EAFP-Med incorporates a prompt-driven dynamic parameter update mechanism, empowering it to extract cross-domain multi-scale lesion features from medical images of diverse modalities adaptively. This exceptional flexibility liberates it from the constraints of any particular imaging technique, fostering great adaptability. Furthermore, EAFP-Med can also serve as a feature preprocessing module connected to any model front-end to enhance the lesion features in input images. Moreover, we propose a novel adaptive disease detection model named EAFP-Med ST, which utilizes the Swin Transformer V2 – Tiny (SwinV2-T) as its backbone and connects it to EAFP-Med. We have compared our method to nine state-of-the-art methods. Experimental results show that the overall accuracy of EAFP Med ST on chest X-ray, brain magnetic resonance imaging, and skin image datasets is 98.47%, 97.60%, and 99.06%, respectively, superior to all the compared state-of-the-art methods.
Xiang Li 0089, Long Lan, Husam Lahza, Shaowu Yang, Shuihua Wang, Wenjing Yang 0002, Hengzhu Liu, Yudong Zhang 0001
Expert Syst. Appl.6
2024 Credit assignment for trained neural networks based on Koopman operator theory
Changyuan Zhao, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002, Zhengbin Pang
Frontiers Comput. Sci.5
2024 Does Confusion Really Hurt Novel Class Discovery?
Haoang Chi, Wenjing Yang 0002, Feng Liu 0003, Long Lan, Bo Han 0003
Int. J. Comput. Vis.2
2024 Qualitative and Quantitative Model Checking Against Recurrent Neural Networks
Wanwei Liu, Fu Song, Bai Xue 0001, Wenjing Yang 0002, Ji Wang 0001, Zhengbin Pang
J. Comput. Sci. Technol.5
2024 Tackling Noisy Labels With Network Parameter Additive Decomposition
abstract
Given data with noisy labels, over-parameterized deep networks suffer overfitting mislabeled data, resulting in poor generalization. The memorization effect of deep networks shows that although the networks have the ability to memorize all noisy data, they would first memorize clean training data, and then gradually memorize mislabeled training data. A simple and effective method that exploits the memorization effect to combat noisy labels is early stopping. However, early stopping cannot distinguish the memorization of clean data and mislabeled data, resulting in the network still inevitably overfitting mislabeled data in the early training stage. In this paper, to decouple the memorization of clean data and mislabeled data, and further reduce the side effect of mislabeled data, we perform additive decomposition on network parameters. Namely, all parameters are additively decomposed into two groups, i.e., parameters w are decomposed as w=σ+γ. Afterward, the parameters σ are considered to memorize clean data, while the parameters γ are considered to memorize mislabeled data. Benefiting from the memorization effect, the updates of the parameters σ are encouraged to fully memorize clean data in early training, and then discouraged with the increase of training epochs to reduce interference of mislabeled data. The updates of the parameters γ are the opposite. In testing, only the parameters σ are employed to enhance generalization. Extensive experiments on both simulated and real-world benchmarks confirm the superior performance of our method.
Xiaobo Xia, Long Lan, Xinghao Wu, Jun Yu 0001, Wenjing Yang 0002, Bo Han 0003, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Verifying safety of neural networks from topological perspectives
Dejin Ren, Bai Xue 0001, Ji Wang 0001, Wenjing Yang 0002, Wanwei Liu
Sci. Comput. Program.5
2024 Out-of-Distribution Generalization With Causal Feature Separation
abstract
Driven by empirical risk minimization, machine learning algorithm tends to exploit subtle statistical correlations existing in the training environment for prediction, while the spurious correlations are unstable across environments, leading to poor generalization performance. Accordingly, the problem of the Out-of-distribution (OOD) generalization aims to exploit an invariant/stable relationship between features and outcomes that generalizes well on all possible environments. To address the spurious correlation induced by the selection bias, in this article, we propose a novel Clique-based Causal Feature Separation (CCFS) algorithm by explicitly incorporating the causal structure to identify causal features of outcome for OOD generalization. Specifically, the proposed CCFS algorithm identifies the largest clique in the learned causal skeleton. Theoretically, we guarantee that either the largest clique or the rest of the causal skeleton is exactly the set of all causal features of the outcome. Finally, we separate the causal features from the non-causal ones with a sample-reweighting decorrelator for OOD prediction. Extensive experiments validate the effectiveness of the proposed CCFS method on both causal feature identification and OOD generalization tasks.
Haotian Wang 0001, Kun Kuang 0001, Long Lan, Zige Wang, Wanrong Huang, Fei Wu 0001, Wenjing Yang 0002
IEEE Trans. Knowl. Data Eng.7
2024 Multiview Deep Anomaly Detection: A Systematic Exploration
abstract
Anomaly detection (AD), which models a given normal class and distinguishes it from the rest of abnormal classes, has been a long-standing topic with ubiquitous applications. As modern scenarios often deal with massive high-dimensional complex data spawned by multiple sources, it is natural to consider AD from the perspective of multiview deep learning. However, it has not been formally discussed by the literature and remains underexplored. Motivated by this blank, this article makes fourfold contributions: First, to the best of our knowledge, this is the first work that formally identifies and formulates the multiview deep AD problem. Second, we take recent advances in relevant areas into account and systematically devise various baseline solutions, which lays the foundation for multiview deep AD research. Third, to remedy the problem that limited benchmark datasets are available for multiview deep AD, we extensively collect the existing public data and process them into more than 30 multiview benchmark datasets via multiple means, so as to provide a better evaluation platform for multiview deep AD. Finally, by comprehensively evaluating the devised solutions on different types of multiview deep AD benchmark datasets, we conduct a thorough analysis on the effectiveness of the designed baselines and hopefully provide other researchers with beneficial guidance and insight into the new multiview deep AD topic.
Siqi Wang 0001, Jiyuan Liu 0003, Xinwang Liu 0002, Sihang Zhou 0001, En Zhu, Yuexiang Yang, Jianping Yin, Wenjing Yang 0002
IEEE Trans. Neural Networks Learn. Syst.9
2024 Boosting Few-shot Object Detection with Discriminative Representation and Class Margin
abstract
Classifying and accurately locating a visual category with few annotated training samples in computer vision has motivated the few-shot object detection technique, which exploits transfering the source-domain detection model to the target domain. Under this paradigm, however, such transferred source-domain detection model usually encounters difficulty in the classification of the target domain because of the low data diversity of novel training samples. To combat this, we present a simple yet effective few-shot detector, Transferable RCNN. To transfer general knowledge learned from data-abundant base classes to data-scarce novel classes, we propose a weight transfer strategy to promote model transferability and an attention-based feature enhancement mechanism to learn more robust object proposal feature representations. Further, we ensure strong discrimination by optimizing the contrastive objectives of feature maps via a supervised spatial contrastive loss. Meanwhile, we introduce an angle-guided additive margin classifier to augment instance-level inter-class difference and intra-class compactness, which is beneficial for improving the discriminative power of the few-shot classification head under a few supervisions. Our proposed framework outperforms the current works in various settings of PASCAL VOC and MSCOCO datasets; this demonstrates the effectiveness and generalization ability.
Shaowu Yang, Wenjing Yang 0002, Dian-xi Shi, Xuehui Li
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Enhanced Dcf Tracker Regularized by Reliable Sample Construction
abstract
Discriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance during template updating and propose an enhanced DCF tracking method regularized by a novel sparse representation based reliable sample construction term, called enhanced sparse correlation filter (ESCF). Specifically, the reconstructed reliable samples are the sparse representation of circulant shifted samples of unfiltered input samples, in which the target will approach the center to preserve target visual cues into the template when using the cosine window. Besides, we jointly perform template learning and reliable sample construction into a unified learning paradigm to benefit from each other, which further can be carried out in the frequency domain without incurring excessive time cost by skillful decomposition. Experiments on several popular visual tracking datasets verify the efficacy of ESCF and show that ESCF performs favorably against several well-established representative counterparts.
Mingyu Cao, Mengzhu Wang, Long Lan, Wenjing Yang 0002, Huibin Tan
ICASSP5
2023 Progressive Perception Learning for Distribution Modulation in Siamese Tracking
abstract
We explore an innovative view on distribution modulation to boost Siamese trackers. Specially, we observed two cases of possible distribution inconsistency in Siamese tracking: 1) Two branches with different sizes may be in different distribution ranges after a shared backbone (including BN layers). 2) The background data may affect the total feature distribution of the search branch. To address these issues, we proposed a plug-and-play component named Progressive Perception Learning Module (P2LM) to modulate the distribution using three feature normalization blocks successively, i.e., Self-Aware Block (SAB), Target-Aware Block (TAB), and Region-Aware Block (RAB). SAB regulates the distribution of each branch independently for the first issue. TAB uses the target information to guide the distribution adjustments of the two branches. RAB divides the search image into foreground and background with a region mask and normalizes them separately to filter the background distractors for robust tracking. TAB and RAB synergistically alleviate the distribution shifts caused by environmental variance. Experiments on OTB100, UAV123, LaSOT, and GOT-10k verify the compelling effects of our module.
Xianchen Zhou, Mingyu Cao, Mengzhu Wang, Guangjie Gao, Wenjing Yang 0002, Huibin Tan
ICASSP6
2023 Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking
abstract
Tensor decomposition and reconstruction attention is a promising global context learning approach because it can remain efficient while avoiding feature compression. To exploit its potential even further in visual tracking, we redesign a 3D tensor modeling paradigm, namely tensor Decomposition, Interaction, Reconstruction attention (DIR), respectively corresponding to three function components, Tensor Decomposition Module (TDM), Tensor Interaction Module (TIM) and Context Reconstruction Module (CRM). Specifically, TDM decomposes a 3D tensor feature into rank-1 context fragments in different dimension views. The ingenuity here lies in the introduction of Circular Convolution for processing features at arbitrary scales and channel-sharing segments to enhance the interaction of the two branches in the Siamese network architecture. TIM obtains the tensor planes of each dimension by the Cross-Similarity operation of rank-1 tensors and fused cubic features, which brings more interactions between all feature dimensions. CRM reconstructs 3D context representations with the outputs of the above modules. In experiments, DIR is embedded into the tracker to verify its effectiveness.
Huibin Tan, Mingyu Cao, Mengzhu Wang, Wenjing Yang 0002
ICASSP6
2023 Domain Specified Optimization for Deployment Authorization
abstract
This paper explores Deployment Authorization (DPA) as a means of restricting the generalization capabilities of vision models on certain domains to protect intellectual property. Nevertheless, the current advancements in DPA are predominantly confined to fully supervised settings. Such settings require the accessibility of annotated images from any unauthorized domain, rendering the DPA approaches impractical for real-world applications due to its exorbitant costs.To address this issue, we propose Source-Only Deployment Authorization (SDPA), which assumes that only authorized domains are accessible during training phases, and the model’s performance on unauthorized domains must be suppressed in inference stages. Drawing inspiration from distributional robust statistics, we present a lightweight method called Domain-Specified Optimization (DSO) for SDPA that degrades the model’s generalization over a divergence ball. DSO comes with theoretical guarantees on the convergence property and its authorization performance. As a complementary of SDPA, we also propose Target-Combined Deployment Authorization (TPDA), where unauthorized domains are partially accessible, and simplify the DSO method to a perturbation operation on the pseudo predictions, referred to as Target-Dependent Domain-Specified Optimization (TDSO). We demonstrate the effectiveness of our proposed DSO and TDSO methods through extensive experiments on six image benchmarks, achieving dominant performance on both SDPA and TDPA settings.
Haotian Wang 0001, Haoang Chi, Wenjing Yang 0002, Mingyang Geng, Long Lan, Jing Zhang 0037, Dacheng Tao
ICCV3
2023 MagicFusion: Boosting Text-to-Image Generation Performance by Fusing Diffusion Models
abstract
The advent of open-source AI communities has produced a cornucopia of powerful text-guided diffusion models that are trained on various datasets. While few explorations have been conducted on ensembling such models to combine their strengths. In this work, we propose a simple yet effective method called Saliency-aware Noise Blending (SNB) that can empower the fused text-guided diffusion models to achieve more controllable generation. Specifically, we experimentally find that the responses of classifier-free guidance are highly related to the saliency of generated images. Thus we propose to trust different models in their areas of expertise by blending the predicted noises of two diffusion models in a saliency-aware manner. SNB is training-free and can be completed within a DDIM sampling process. Additionally, it can automatically align the semantics of two noise spaces without requiring additional annotations such as masks. Extensive experiments show the impressive effectiveness of SNB in various applications. The project page is available at https://magicfusion.github.io/.
Heliang Zheng, Long Lan, Wenjing Yang 0002
ICCV5
2023 A Geometrical Characterization on Feature Density of Image Datasets
abstract
Recently, the interpretability and verification of deep learning have attracted enormous attention from both academic and industrial communities, aiming to gain users’ trust and ease their concerns. To guide learning procedures or data operations carried out in a more interpretable way, in this paper, we put a similar perspective on image datasets, the inputs of deep learning. Based on manifold learning, we work out an interpretable geometrical characterization on the curvity of manifolds to depict the feature density of datasets, which is represented with the ratio of the Euclidean distance and the geodesic distance. It is a noteworthy characteristic of image datasets and we take the dataset compression and enhancement problems as application instances via sample credit assignment with the geometrical information. Experiments on typical image datasets have justified the effectiveness and enormous prospect of the presented geometrical characteristic.
Changyuan Zhao, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002
ICME5
2023 Memory-based Exploration-value Evaluation Model for Visual Navigation
abstract
We propose a hierarchical visual navigation solution, called Memory-based Exploration-value Evaluation Model (MEEM), to improve the agent's navigation performance. MEEM employs a hierarchical policy to tackle the challenge of sparse rewards, holds an episodic memory to store the historical information of the agent, and applies an Exploration-value Evaluation Model to calculate an exploration-value for action planning at each location in the observable area. We experimentally verify MEEM by navigation performance comparison on two datasets including the grid-map dataset and the 3D scenes Gibson dataset, where our approach achieves state-of-the-art performance on both. Specifically, the overall success rate of MEEM is 95% on the grid-map dataset while the best competitor reaches 68% only. As for the Gibson dataset, the success rate of ours and the best competitor SemExp are 69.8% and 54.4%, respectively. Ablation analysis on the tile-map dataset indicates that all three components of MEEM have positive effects.
Yongquan Feng, Minglong Li, Ruochun Jin, Shaowu Yang, Wenjing Yang 0002
ICRA7
2023 GANet: Goal Area Network for Motion Forecasting
abstract
Predicting the future motion of road participants is crucial for autonomous driving but is extremely challenging due to staggering motion uncertainty. Recently, most motion forecasting methods resort to the goal-based strategy, i.e., predicting endpoints of motion trajectories as conditions to regress the entire trajectories, so that the search space of solution can be reduced. However, accurate goal coordinates are hard to predict and evaluate. In addition, the point representation of the destination limits the utilization of a rich road context, leading to inaccurate prediction results in many cases. Goal area, i.e., the possible destination area, rather than goal coordinate, could provide a more soft constraint for searching potential trajectories by involving more tolerance and guidance. In view of this, we propose a new goal area-based framework, named Goal Area Network (GANet), for motion forecasting, which models goal areas as preconditions for trajectory prediction, performing more robustly and accurately. Specifically, we propose a GoICrop (Goal Area of Interest) operator to effectively aggregate semantic lane features in goal areas and model actors' future interactions as feedback, which benefits a lot for future trajectory estimations. GANet ranks the 1st on the leaderboard of Argoverse Challenge among all public literature (till the paper submission). Code will be available at https://github.com/kingwmk/GANet.
Mingkun Wang, Xinge Zhu, Changqian Yu, Wei Li 0111, Yuexin Ma, Ruochun Jin, Xiaoguang Ren, Dongchun Ren, Wenjing Yang 0002
ICRA10
2023 Task2Morph: Differentiable Task-Inspired Framework for Contact-Aware Robot Design
abstract
Optimizing the morphologies and the controllers that adapt to various tasks is a critical issue in the field of robot design, aka. embodied intelligence. Previous works typically model it as a joint optimization problem and use search-based methods to find the optimal solution in the morphology space. However, they ignore the implicit knowledge of task-to-morphology mapping which can directly inspire robot design. For example, flipping heavier boxes tends to require more muscular robot arms. This paper proposes a novel and general differentiable task-inspired framework for contact-aware robot design called Task2Morph. We abstract task features highly related to task performance and use them to build a task-to-morphology mapping. Further, we embed the mapping into a differentiable robot design process, where the gradient information is leveraged for both the mapping learning and the whole optimization. The experiments are conducted on three scenarios, and the results validate that Task2Morph outperforms DiffHand, which lacks a task-inspired morphology module, in terms of efficiency and effectiveness.
Yishuai Cai, Shaowu Yang, Minglong Li, Xinglin Chen, Yunxin Mao, Xiaodong Yi 0002, Wenjing Yang 0002
IROS7
2023 Evolving Physical Instinct for Morphology and Control Co-Adaption
abstract
The capability of a robot to perform tasks depends not only on precise motion control, but also on a well-suited body morphology. Adapting both morphology and control of robots to improve their task performance has been a widely studied and long-standing issue. While the bio-inspired bi-level optimization framework has gained popularity in recent years, it suffers from high computation complexity due to the time-consuming and inefficient learning process for each morphology. In fact, in nature, besides the adaptive morphology and the intelligent brain, animals also possess an important gift, which is physical instinct. These instincts allow animals to respond quickly to their surroundings in the neonatal period, facilitating skills acquisition. Inspired by this, we propose an evolvable instinct controller to enhance the morphology-control co-adaption. The instinct controller suggests rough motion inclinations, which require minimal domain knowledge and entail less sophisticated design. Its purpose is to assist the main controller in learning fine-grained and robust control efficiently. We implemented this idea in the context of legged locomotion and designed the instinct controller using phase-based FSMs. We propose the instinct-based co-adaption algorithm and construct GPU parallel simulation experiments on different morphology prototypes. The results indicate that combining the co-adaption process with instinct evolution leads to the development of superior morphologies and robust controllers compared with the conventional co-adaption approach, with minimal additional time cost.
Xinglin Chen, Minglong Li, Yishuai Cai, Zhuoer Wen, Zhongxuan Cai, Wenjing Yang 0002
IROS7
2023 Treatment Effect Estimation with Adjustment Feature Selection
abstract
In causal inference, it is common to select a subset of observed covariates, named the adjustment features, to be adjusted for estimating the treatment effect. For real-world applications, the abundant covariates are usually observed, which contain extra variables partially correlating to the treatment (treatment-only variables, e.g., instrumental variables) or the outcome (outcome-only variables, e.g., precision variables) besides the confounders (variables that affect both the treatment and outcome). In principle, unbiased treatment effect estimation is achieved once the adjustment features contain all the confounders. However, the performance of empirical estimations varies a lot with different extra variables. To solve this issue, variable separation/selection for treatment effect estimation has received growing attention when the extra variables contain instrumental variables and precision variables.
Haotian Wang 0001, Kun Kuang 0001, Haoang Chi, Longqi Yang 0002, Mingyang Geng, Wanrong Huang, Wenjing Yang 0002
KDD7
2023 Null-text Guidance in Diffusion Models is Secretly a Cartoon-style Creator
abstract
Classifier-free guidance is an effective sampling technique in diffusion models that has been widely adopted. The main idea is to extrapolate the model in the direction of text guidance and away from null-text guidance. In this paper, we demonstrate that null-text guidance in diffusion models is secretly a cartoon-style creator, i.e., the generated images can be efficiently transformed into cartoons by simply perturbing the null-text guidance. Specifically, we proposed two disturbance methods, i.e., Rollback disturbance (Back-D) and Image disturbance (Image-D), to construct misalignment between the noisy images used for predicting null-text guidance and text guidance (subsequently referred to as null-text noisy image and text noisy imageb respectively) in the sampling process. Back-D achieves cartoonization by altering the noisb level of the null-text noisy image via replacing xt with xl + Δ t. Image-D, alternatively, produces high-fidelity, diverse cartoons by defining xt as a clean input image, which further improves the incorporation of finer image details. Through comprehensive experiments, we delved into the principle of noise disturbing for null-text and uncovered that the efficacy of disturbance depends on the correlation between the null-text noisy image and the source image. Moreover, the proposed methods, which can generate cartoon images and cartoonize specific ones, are training-free and easily integrated as a plug-and-play component in any classifier-free guided diffusion model. The project page is available at https://nulltextforcartoon.github.io/.
Heliang Zheng, Long Lan, Wanrong Huang, Wenjing Yang 0002
ACM Multimedia6
2023 SODA: Robust Training of Test-Time Data Adaptors
abstract
Adapting models deployed to test distributions can mitigate the performance degradation caused by distribution shifts. However, privacy concerns may render model parameters inaccessible. One promising approach involves utilizing zeroth-order optimization (ZOO) to train a data adaptor to adapt the test data to fit the deployed models. Nevertheless, the data adaptor trained with ZOO typically brings restricted improvements due to the potential corruption of data features caused by the data adaptor. To address this issue, we revisit ZOO in the context of test-time data adaptation. We find that the issue directly stems from the unreliable estimation of the gradients used to optimize the data adaptor, which is inherently due to the unreliable nature of the pseudo-labels assigned to the test data. Based on this observation, we propose pseudo-label-robust data adaptation (SODA) to improve the performance of data adaptation. Specifically, SODA leverages high-confidence predicted labels as reliable labels to optimize the data adaptor with ZOO for label prediction. For data with low-confidence predictions, SODA encourages the adaptor to preserve data information to mitigate data corruption. Empirical results indicate that SODA can significantly enhance the performance of deployed models in the presence of distribution shifts without requiring access to model parameters.
Zige Wang, Yonggang Zhang 0003, Zhen Fang 0001, Long Lan, Wenjing Yang 0002, Bo Han 0003
NeurIPS5
2023 Safety Verification for Neural Networks Based on Set-Boundary Analysis
Dejin Ren, Wanwei Liu, Ji Wang 0001, Wenjing Yang 0002, Bai Xue 0001
TASE5
2023 Self-aware circular response-guided attention for robust siamese tracking
Huibin Tan, Mengzhu Wang, Tianyi Liang 0001, Yuhua Tang, Long Lan, Wenjing Yang 0002
Appl. Intell.7
2023 Towards robust neural networks via a global and monotonically decreasing robustness training strategy
abstract
Robustness of deep neural networks (DNNs) has caused great concerns in the academic and industrial communities, especially in safety-critical domains. Instead of verifying whether the robustness property holds or not in certain neural networks, this paper focuses on training robust neural networks with respect to given perturbations. State-of-the-art training methods, interval bound propagation (IBP) and CROWN-IBP, perform well with respect to small perturbations, but their performance declines significantly in large perturbation cases, which is termed “drawdown risk” in this paper. Specifically, drawdown risk refers to the phenomenon that IBP-family training methods cannot provide expected robust neural networks in larger perturbation cases, as in smaller perturbation cases. To alleviate the unexpected drawdown risk, we propose a global and monotonically decreasing robustness training strategy that takes multiple perturbations into account during each training epoch (global robustness training), and the corresponding robustness losses are combined with monotonically decreasing weights (monotonically decreasing robustness training). With experimental demonstrations, our presented strategy maintains performance on small perturbations and the drawdown risk on large perturbations is alleviated to a great extent. It is also noteworthy that our training method achieves higher model accuracy than the original training methods, which means that our presented training strategy gives more balanced consideration to robustness and accuracy.
Taoran Wu, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002, Ji Wang 0001, Zhengbin Pang
Frontiers Inf. Technol. Electron. Eng.5
2023 Recent Advances for Quantum Neural Networks in Generative Learning
abstract
Quantum computers are next-generation devices that hold promise to perform calculations beyond the reach of classical computers. A leading method towards achieving this goal is through quantum machine learning, especially quantum generative learning. Due to the intrinsic probabilistic nature of quantum mechanics, it is reasonable to postulate that quantum generative learning models (QGLMs) may surpass their classical counterparts. As such, QGLMs are receiving growing attention from the quantum physics and computer science communities, where various QGLMs that can be efficiently implemented on near-term quantum machines with potential computational advantages are proposed. In this paper, we review the current progress of QGLMs from the perspective of machine learning. Particularly, we interpret these QGLMs, covering quantum circuit Born machines, quantum generative adversarial networks, quantum Boltzmann machines, and quantum variational autoencoders, as the quantum extension of classical generative learning models. In this context, we explore their intrinsic relations and their fundamental differences. We further summarize the potential applications of QGLMs in both conventional machine learning tasks and quantum physics. Last, we discuss the challenges and further research directions for QGLMs.
Jinkai Tian, Shanshan Zhao 0001, Qing Liu 0027, Kaining Zhang, Wanrong Huang, Xingyao Wu, Min-Hsiu Hsieh, Tongliang Liu, Wenjing Yang 0002, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.13
2022 Meta Discovery: Learning to Discover Novel Classes given Very Limited Data
Haoang Chi, Feng Liu 0003, Wenjing Yang 0002, Long Lan, Tongliang Liu, Bo Han 0003, Gang Niu 0001, Mingyuan Zhou, Masashi Sugiyama
ICLR3
2022 Placement Optimization for UAV-Enabled Wireless Networks with Multi-Hop Backhauls in Urban Environments
abstract
In surveillance or search scenarios, exploiting unmanned aerial vehicles (UAVs) as relays to provide wireless data access for task-oriented ground robots (GRs) with remote base station have emerged as a promising application. This paper considers a UAV-enabled wireless network, where communication links could be line-of-sight (LoS) and non-line-of-sight (NLoS) due to obstacles in urban environments. Existing works typically adopted the free-space path loss model or the statistical channel model, which either ignored the impact of obstacles or assumed uniformly distributed obstacles and therefore might fail in practical NLoS scenarios. In this paper, taking the information of randomly distributed obstacles in environments into consideration, we aim to optimize the placement for the UAV-enabled multi-hop network to transfer more data collected by GRs and minimize the time delay in data transmission while satisfying the required communication quality. By reconstructing this complex non-convex optimization problem into two subprob-lems and solving them alternatively, we propose the multi-hop UAVs placement (mUP) method to get the solution, which contains the air-to-ground network formation (ATG-NF) algorithm and the communication quality-aware UAV placement (CQA-UP) algorithm. Simulation results show that in four types of typical urban environments or with different numbers of UAVs, the proposed mUP method achieves substantial performance gains in terms of communication quality and task performance compared to other placement approaches based on statistical channel models. We further discuss the robustness of the mUP method towards terrain measurement error.
Sining Yang, Dian-xi Shi, Yingxuan Peng, Shaowu Yang, Bo Zhang 0007, Wenjing Yang 0002
IPSN6
2022 Estimating Individualized Causal Effect with Confounded Instruments
abstract
Learning individualized causal effect (ICE) plays a vital role in various fields of big data analysis, ranging from fine-grained policy evaluation to personalized treatment development. However, the presence of unmeasured confounders increases the difficulty of estimating ICE in real-world scenarios. A wide range of methods have been proposed to address the unmeasured confounders with the aid of instrument variable (IV), which sources from the treatment randomization. The performance of these methods relies on the well-predefined IVs that satisfy the unconfounded instruments assumption (i.e., the IVs are independent with the unmeasured confounders given observed covariates), which is untestable and leads to finding a valid IV becomes an art rather than science. In this paper, we focus on estimating the ICE with confounded instruments that violate the unconfounded instruments assumption. By considering the conditional independence between the set of confounded instruments and the outcome variable, we propose a novel method, named CVAE-IV, to generate a substitute of the unmeasured confounder with a conditional variational autoencoder. Our theoretical analysis guarantees that the generated confounder substitute will identify unbiased ICE. Extensive experiments on bias demand prediction and Mendelian randomization analysis verify the effectiveness of our method.
Haotian Wang 0001, Wenjing Yang 0002, Longqi Yang 0002, Anpeng Wu, Fei Wu 0001, Kun Kuang 0001
KDD2
2022 RepSSRN: The Structural Reparameterization Applied to SSRN for Hyperspectral Image Classification
abstract
Recently, deep learning has been widely used in hyperspectral image classification. Reparameterized network as a novel deep learning method achieves competitive performance compared to other methods. Hyperspectral images have the defect of a large amount of data, and reparameterization can improve network performance without increasing network complexity because of the structure fusion. However, according to our survey there is no relevant application in hyperspectral classification. In this paper, we propose a network using reparameterized approach, which is applied in the field of hyperspectral image classification firstly. This network reparameterizes the Spectral-Spatial Residual Network (SSRN), abbreviated as RepSSRN. Experiments in three classical datasets show that RepSSRN has improved the performance of SSRN.
Yuqian Wu, Lulu Shi, Wenjing Yang 0002
IEEE Geosci. Remote. Sens. Lett.5
2022 Heterogeneous Pseudo-Supervised Learning for Few-shot Person Re-Identification
Long Lan, Wenjing Yang 0002
Neural Networks5
2021 BT Expansion: a Sound and Complete Algorithm for Behavior Planning of Intelligent Robots with Behavior Trees
abstract
Behavior Trees (BTs) have attracted much attention in the robotics field in recent years, which generalize existing control architectures and bring unique advantages for building robot systems. Automated synthesis of BTs can reduce human workload and build behavior models for complex tasks beyond the ability of human design, but theoretical studies are almost missing in existing methods because it is difficult to conduct formal analysis with the classic BT representations. As a result, they may fail in tasks that are actually solvable. This paper proposes BT expansion, an automated planning approach to building intelligent robot behaviors with BTs, and proves the soundness and completeness through the state-space formulation of BTs. The advantages of blended reactive planning and acting are formally discussed through the region of attraction of BTs, by which robots with BT expansion are robust to any resolvable external disturbances. Experiments with a mobile manipulator and test sets are simulated to validate the effectiveness and efficiency, where the proposed algorithm surpasses the baseline by virtue of its soundness and completeness. To the best of our knowledge, it is the first time to leverage the state-space formulation to synthesize BTs with a complete theoretical basis.
Zhongxuan Cai, Minglong Li, Wanrong Huang, Wenjing Yang 0002
AAAI4
2021 Dec-SGTS: Decentralized Sub-Goal Tree Search for Multi-Agent Coordination
abstract
Multi-agent coordination tends to benefit from efficient communication, where cooperation often happens based on exchanging information about what the agents intend to do, i.e. intention sharing. It becomes a key problem to model the intention by some proper abstraction. Currently, it is either too coarse such as final goals or too fined as primitive steps, which is inefficient due to the lack of modularity and semantics. In this paper, we design a novel multi-agent coordination protocol based on subgoal intentions, defined as the probability distribution over feasible subgoal sequences. The subgoal intentions encode macro-action behaviors with modularity so as to facilitate joint decision making at higher abstraction. Built over the proposed protocol, we present Dec-SGTS (Decentralized Sub-Goal Tree Search) to solve decentralized online multi-agent planning hierarchically and efficiently. Each agent runs Dec-SGTS asynchronously by iteratively performing three phases including local sub-goal tree search, local subgoal intention update and global subgoal intention sharing. We conduct the experiments on courier dispatching problem, and the results show that Dec-SGTS achieves much better reward while enjoying a significant reduction of planning time and communication cost compared with Dec-MCTS (Decentralized Monte Carlo Tree Search).
Minglong Li, Zhongxuan Cai, Wenjing Yang 0002, Lixia Wu, Ji Wang 0001
AAAI3
2021 Conditional Variational Capsule Network for Open Set Recognition
abstract
In open set recognition, a classifier has to detect unknown classes that are not known at training time. In order to recognize new categories, the classifier has to project the input samples of known classes in very compact and separated regions of the features space for discriminating samples of unknown classes. Recently proposed Capsule Networks have shown to outperform alternatives in many fields, particularly in image recognition, however they have not been fully applied yet to open-set recognition. In capsule networks, scalar neurons are replaced by capsule vectors or matrices, whose entries represent different proper-ties of objects. In our proposal, during training, capsules features of the same known class are encouraged to match a pre-defined gaussian, one for each class. To this end, we use the variational autoencoder framework, with a set of gaussian priors as the approximation for the posterior distribution. In this way, we are able to control the compactness of the features of the same class around the center of the gaussians, thus controlling the ability of the classifier in detecting samples from unknown classes. We conducted several experiments and ablation of our model, obtaining state of the art results on different datasets in the open set recognition and unknown detection tasks.
Yunrui Guo, Guglielmo Camporese, Wenjing Yang 0002, Alessandro Sperduti, Lamberto Ballan
ICCV3
2021 TOHAN: A One-step Approach towards Few-shot Hypothesis Adaptation
abstract
In few-shot domain adaptation (FDA), classifiers for the target domain are trained with \emph{accessible} labeled data in the source domain (SD) and few labeled data in the target domain (TD). However, data usually contain private information in the current era, e.g., data distributed on personal phones. Thus, the private data will be leaked if we directly access data in SD to train a target-domain classifier (required by FDA methods). In this paper, to prevent privacy leakage in SD, we consider a very challenging problem setting, where the classifier for the TD has to be trained using few labeled target data and a well-trained SD classifier, named few-shot hypothesis adaptation (FHA). In FHA, we cannot access data in SD, as a result, the private information in SD will be protected well. To this end, we propose a target-oriented hypothesis adaptation network (TOHAN) to solve the FHA problem, where we generate highly-compatible unlabeled data (i.e., an intermediate domain) to help train a target-domain classifier. TOHAN maintains two deep networks simultaneously, in which one focuses on learning an intermediate domain and the other takes care of the intermediate-to-target distributional adaptation and the target-risk minimization. Experimental results show that TOHAN outperforms competitive baselines significantly.
Haoang Chi, Feng Liu 0003, Wenjing Yang 0002, Long Lan, Tongliang Liu, Bo Han 0003, William Kwok-Wai Cheung, James T. Kwok
NeurIPS3
2021 Graph Adversarial Self-Supervised Learning
abstract
This paper studies a long-standing problem of learning the representations of a whole graph without human supervision. The recent self-supervised learning methods train models to be invariant to the transformations (views) of the inputs. However, designing these views requires the experience of human experts. Inspired by adversarial training, we propose an adversarial self-supervised learning (\texttt{GASSL}) framework for learning unsupervised representations of graph data without any handcrafted views. \texttt{GASSL} automatically generates challenging views by adding perturbations to the input and are adversarially trained with respect to the encoder. Our method optimizes the min-max problem and utilizes a gradient accumulation strategy to accelerate the training process. Experimental on ten graph classification datasets show that the proposed approach is superior to state-of-the-art self-supervised learning baselines, which are competitive with supervised models.
Longqi Yang 0002, Wenjing Yang 0002
NeurIPS3
2021 ANF: Attention-Based Noise Filtering Strategy for Unsupervised Few-Shot Classification
Guangsen Ni, Wenjing Yang 0002, Long Lan
PRICAI (3)5
2021 Role-based attention in deep reinforcement learning for games
abstract
Abstract Reinforcement learning method that learns while interacting with the environment, relies heavily on the concept of state as the input to the policy and value function. In the task, the view of agent contains a lot of information, and it is difficult for the agent to learn to ignore the irrelevant information and focus on the key information. Inspired by recent work in attention models for computer vision, we present a role‐based attention model for reinforcement learning. The proposed model uses convolutional neural networks to generate soft attention maps, adding crucial role information in the task, forcing the agent to focus on important features and distinguish task‐related information. To validate the performance in complex problems, the proposed approach is evaluated in a challenging scenario, Football Academy in Google Research Football Environment, a newly released reinforcement learning environment with physics‐based three‐dimensional simulator. The experimental results demonstrate that agents using role‐based attention mechanism can perform better in football games.
Dong Yang 0010, Wenjing Yang 0002, Minglong Li, Qiong Yang
Comput. Animat. Virtual Worlds2
2021 A robust quadruple adaptation network in few-shot scenarios
Haoang Chi, Shengang Li, Wenjing Yang 0002, Long Lan
Knowl. Based Syst.3
2020 Robust Normalized Squares Maximization for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) attempts to transfer specific knowledge from one domain with labeled data to another domain without labels. Recently, maximum squares loss has been proposed to tackle UDA problem but it does not consider the prediction diversity which has proven beneficial to UDA. In this paper, we propose a novel normalized squares maximization (NSM) loss in which the maximum squares is normalized by the sum of squares of class sizes. The normalization term enforces the class sizes of predictions to be balanced to explicitly increase the diversity. Theoretical analysis shows that the optimal solution to NSM is one-hot vectors with balanced class sizes, i.e., NSM encourages both discriminate and diverse predictions. We further propose a robust variant of NSM, RNSM, by replacing the square loss with L2,1-norm to reduce the influence of outliers and noises. Experiments of cross-domain image classification on two benchmark datasets illustrate the effectiveness of both NSM and RNSM. RNSM achieves promising performance compared to state-of-the-art methods. The code is available at https://github.com/wj-zhang/NSM.
Wenju Zhang, Xiang Zhang 0008, Qing Liao 0001, Wenjing Yang 0002, Long Lan, Zhigang Luo
CIKM4
2020 Pairwise Similarity Regularization for Adversarial Domain Adaptation
abstract
Domain adaptation aims at learning a predictive model that can generalize to a new target domain different from the source (training) domain. To mitigate the domain gap, adversarial training has been developed to learn domain invariant representations. State-of-the-art methods further make use of pseudo labels generated by the source domain classifier to match conditional feature distributions between the source and target domains. However, if the target domain is more complex than the source domain, the pseudo labels are unreliable to characterize the class-conditional structure of the target domain data, undermining prediction performance. To resolve this issue, we propose a Pairwise Similarity Regularization (PSR) approach that exploits cluster structures of the target domain data and minimizes the divergence between the pairwise similarity of clustering partition and that of pseudo predictions. Therefore, PSR guarantees that two target instances in the same cluster have the same class prediction and thus eliminate the negative effect of unreliable pseudo labels. Extensive experimental results show that our PSR method significantly boosts the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, PSR achieves a remarkable improvement of more than 5% over the state-of-the-art on several hard-to-transfer tasks.
Haotian Wang 0001, Wenjing Yang 0002, Ji Wang 0001, Ruxin Wang 0002, Long Lan, Mingyang Geng
ACM Multimedia2
2020 One-shot video-based person re-identification with variance subsampling algorithm
abstract
Abstract Previous works propose the distance‐based sampling for unlabeled datapoints to address the few‐shot person re‐identification task, however, many selected samples may be assigned with wrong labels due to poor feature quality in these works, which negatively affects the learning procedure. In this article, we propose a novel sampling strategy to improve the quality of assigned pseudo‐labels, thus promoting the final performance. To illustrate, we first propose the concept of variance confidence to measure the credibility of pseudo‐labels, then we apply a novel variance subsampling algorithm to improve the accuracy of the selected sample labels. Our method combines distance confidence and variance confidence as a two‐round sampling criterion. Meanwhile, a variation decay strategy is used in our sampling process in combination with the actual distribution of features. We evaluate our approach on two publicly available datasets, MARS and DukeMTMC‐VideoReID, and achieve state‐of‐the‐art one‐shot performance.
Wenjing Yang 0002, Wanrong Huang, Qiong Yang
Comput. Animat. Virtual Worlds2
2019 TMDA: Task-Specific Multi-source Domain Adaptation via Clustering Embedded Adversarial Training
abstract
Beyond classical domain-specific adversarial training, a recently proposed task-specific framework has achieved a great success in single source domain adaptation by utilizing task-specific decision boundaries. However, compared to single-source-single-target setting, multi-source domain adaptation (MDA) shows more powerful capability to handle with most real-life cases. To align target domain with diverse multi-source domains using task-specific decision boundaries, we provide a deep insight of task-specific framework on MDA for the first time. Accordingly, we propose a novel task-specific multi-source domain adaptation method (TMDA) with a clustering embedded adversarial training process. Specifically, the proposed TMDA detects and refines less discriminative target representations through a max-min optimization over two adversarial task-specific classifiers. Moreover, our analysis implies that scattered multi-source representations disturb the adversarial training under the task-specific framework. To tight up the dispersed source representations, we embeds a relationship-based domain clustering into TMDA. Empirical results demonstrate that our TMDA outperforms state-of-the-art methods on toy dataset, sentiment analysis and digit classification.
Haotian Wang 0001, Wenjing Yang 0002
ICDM2
2019 Non-Convex Transfer Subspace Learning for Unsupervised Domain Adaptation
abstract
Transfer subspace learning aims to learn robust subspace for the target domain by leveraging knowledge from the source domain. The traditional methods often adopt the convex norm to approximate the original sparse and low-rank constraints, which make the optimization problem be easily solved. However, such relax approximation leads to the performance deviation of the original non-convex model. In this paper, we propose a novel Non-convex Transfer Subspace Learning~(NTSL) method to provide a tighter approximation to the original sparse and low-rank constraints. Specifically, we design an objective function that leverages the Schatten p-norm and ℓ_2, p-norm to preserve the structure between the source and target domains. With Schatten p-norm, the objective function better approximates the rank minimization problem than the nuclear norm and preserves the structure of domains. Besides, the ℓ_2, p-norm can reduce the effect of noise and improve the robustness to outliers. Meanwhile, we develop an efficient algorithm to solve the non-convex minimization problem. Extensive experimental results on cross-domain tasks show the effectiveness of our proposed method.
Tingjin Luo, Wenjing Yang 0002, Yongjun Zhang 0006, Yuhua Tang
ICME4
2019 Parallel Gym Gazebo: a Scalable Parallel Robot Deep Reinforcement Learning Platform
abstract
Deep reinforcement learning is making advances in robotics with the platforms of realistic environment simulation. However, as shown in this paper, the realistic simulation introduces vast time cost which is the bottleneck of the learning procedure. To solve this problem generally, we propose a parallel reinforcement learning platform which follows the master-slave principle and integrates learning programs with multiple distributedrobot simulators. The platform is intrinsically scalable and requires no modification to existing serially designed learning environments or algorithms. Experimental results demonstrate that our platform significantly accelerates the learning progress of robots, in direct proportion to the parallel scale. The parallelism also brings richer exploration and sampling, enhancing the performance of deep reinforcement learning algorithms compared with existing serial platforms.
Zhongxuan Cai, Minglong Li, Wenjing Yang 0002
ICTAI4
2019 Integrating Decision Sharing with Prediction in Decentralized Planning for Multi-Agent Coordination under Uncertainty
abstract
The performance of decentralized multi-agent systems tends to benefit from information sharing and its effective utilization. However, too much or unnecessary sharing may hinder the performance due to the delay, instability and additional overhead of communications. Aiming to a satisfiable coordination performance, one would prefer the cost of communications as less as possible. In this paper, we propose an approach for improving the sharing utilization by integrating information sharing with prediction in decentralized planning. We present a novel planning algorithm by combining decision sharing and prediction based on decentralized Monte Carlo Tree Search called Dec-MCTS-SP. Each agent grows a search tree guided by the rewards calculated by the joint actions, which can not only be sampled from the shared probability distributions over action sequences, but also be predicted by a sufficiently-accurate and computationally-cheap heuristics-based method. Besides, several policies including sparse and discounted UCT and DIY-bonus are leveraged for performance improvement. We have implemented Dec-MCTS-SP in the case study on multi-agent information gathering under threat and uncertainty, which is formulated as Decentralized Partially Observable Markov Decision Process (Dec-POMDP). The factored belief vectors are integrated into Dec-MCTS-SP to handle the uncertainty. Comparing with the random, auction-based algorithm and Dec-MCTS, the evaluation shows that Dec-MCTS-SP can reduce communication cost significantly while still achieving a surprisingly higher coordination performance.
Minglong Li, Wenjing Yang 0002, Zhongxuan Cai, Shaowu Yang, Ji Wang 0001
IJCAI2
2019 Rademacher dropout: An adaptive dropout for deep neural network via optimizing generalization gap
Haotian Wang 0001, Wenjing Yang 0002, Tingjin Luo, Ji Wang 0001, Yuhua Tang
Neurocomputing2
2019 Cauchy sparse NMF with manifold regularization: A robust method for hyperspectral unmixing
Haotian Wang 0001, Wenjing Yang 0002, Naiyang Guan
Knowl. Based Syst.2
2018 MulAttenRec: A Multi-level Attention-Based Model for Recommendation
Wenjing Yang 0002, Yongjun Zhang 0006, Haotian Wang 0001, Yuhua Tang
ICONIP (2)2
2018 Multi-feature Fusion for Deep Reinforcement Learning: Sequential Control of Mobile Robots
Haotian Wang 0001, Wenjing Yang 0002, Wanrong Huang, Yuhua Tang
ICONIP (7)2
2018 Deep CNN-based Visual Target Tracking System Relying on Monocular Image Sensing
abstract
The one-on-one target tracking problem is important in robot vision. Previous studies mainly focused on locating, depth information and control mechanism. In this study, we construct an autonomously visual tracking system called learn-to-track (LtT) by using a novel approach. This system only depends on a monocular camera. The main component is a deep convolutional neural network called the LtT, which trains a supervised image classifier by using images captured by the monocular camera in the follower robot. By operating merely on two adjacent frames, the network can predict the estimated velocity of the target, i.e., the velocity control for the follower. To verify the effectiveness of the LtT system, we construct a large-scale dataset that supports download l in the simulator, in which the LtT network is trained and the LtT system performance is evaluated. Furthermore, a remarkable tracking performance is achieved.
Yawen Cui, Bo Zhang 0007, Wenjing Yang 0002, Xiaodong Yi 0002, Yuhua Tang
IJCNN3