EDBT 2026 Demo / reviewers in the wild / expert
Wanrong Huang
dblp:184/0874
· DBLP profile ↗
29ranked-venue papers
2as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMG-VL: A Vision-Language Driven Approach for Multi-Person Motion GenerationabstractGenerating realistic and coordinated 3D human motion for multiple individuals within complex environments remains a significant challenge. Existing text-to-motion methods are often ``blind'' to the physical scene, leading to implausible motions, while scene-conditioned (HSI) approaches demand cumbersome full 3D data and largely neglect multi-person dynamics. To address these limitations, we introduce the VL2Motion paradigm and its embodiment, MMG-VL, a hierarchical framework that generates coordinated multi-person motions from the most accessible inputs: a single 2D image and natural language. MMG-VL first employs a Scene-Aware Intent Planner (SAIP) to interpret the visual context and decompose the user's command into a set of spatially-grounded, multi-person action blueprints. Subsequently, a Coordinated Motion Synthesizer (CMS) translates these blueprints into high-fidelity 3D motion sequences. The synergy between these stages is driven by two novel loss functions: a Spatial-Semantic Grounding Loss to ensure the planner's output is grounded in visual reality, and a Coordinated Environmental Realism Loss that enforces physical constraints and coherent group dynamics during synthesis. To facilitate this research, we introduce HumanVL, the first large-scale dataset featuring multi-person activities in multi-room scenes, providing aligned images, text, blueprints, 3D motions, and scene geometry. Extensive experiments demonstrate that MMG-VL significantly outperforms existing methods in generating spatially coherent, physically realistic, and coordinated multi-person motions, paving the way for more scalable and intuitive creation of dynamic virtual worlds. Songyuan Yang, Wanrong Huang, Yinuo Liu, Kedi Zhang, Xihuai He, Shaowu Yang, Huibin Tan |
AAAI | 2 |
| 2026 | Leveraging VLMs for MUDA: Category-specific prompt with multi-modal interactive LoRA
Xihuai He, Xueqiong Li, Wanrong Huang, Hengzhu Liu, Huibin Tan |
Neural Networks | 4 |
| 2025 | MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion ModelsabstractLarge-scale text-to-image diffusion models, (e.g., DALL-E, SDXL) are capable of generating famous persons by simply referring to their names. Is it possible to make such models generate generic identities as simple as the famous ones, e.g., just use a name? In this paper, we explore the existence of a ``Name Space'', where any point in the space corresponds to a specific identity. Fortunately, we find some clues in the feature space spanned by text embedding of celebrities' names. Specifically, we first extract the embeddings of celebrities' names in the Laion5B dataset with the text encoder of diffusion models. Such embeddings are used as supervision to learn an encoder that can predict the name (actually an embedding) of a given face image. We experimentally find that such name embeddings work well in promising the generated image with good identity consistency. Note that like the names of celebrities, our predicted name embeddings are disentangled from the semantics of text inputs, making the original generation capability of text-to-image models well-preserved. Moreover, by simply plugging such name embeddings, all variants (e.g., from Civitai) derived from the same base model (i.e., SDXL) readily become identity-aware text-to-image models. Heliang Zheng, Long Lan, Wanrong Huang, Yuhua Tang |
AAAI | 5 |
| 2025 | Robust CLIP-Guided Deep Thinking: A Two-Stage Optimization Strategy for Enhancing Adversarial Robustness and Reliability in LVLMsabstractLarge Vision-Language models (LVLMs) have demonstrated remarkable performance in a wide range of vision-language tasks as an efficient input/output system. However, the lack of adversarial robustness at the input side and the widespread hallucination phenomenon at the output side significantly undermine user trust in them. Current solutions to the former tend to sacrifice the general performance of LVLMs, while solving the latter requires a large amount of engineering costs. To address these challenges, we propose a two-stage optimization strategy called RCDT (Robust CLIP-guided Deep Thinking), which aims to enhance the adversarial robustness of LVLMs with minimal general performance loss while reducing hallucinations. First, we introduce a constrained adversarial fine-tuning approach for CLIP to limit the general performance loss during the enhancement of robustness. Furthermore, this CLIP is used to think deeply about the output process of LVLMs to reduce hallucinations. Experiments show that RCDT not only reduce general performance loss by more than half while maintaining adversarial robustness compared to the baselines, but also demonstrate good performance in mitigating hallucinations. Yize Sui, Wanrong Huang, Wenjing Yang 0002, Chaofan Zhao, Ji Wang 0001 |
ICASSP | 2 |
| 2025 | Effective and Efficient Time-Varying Counterfactual Prediction with State-Space ModelsabstractTime-varying counterfactual prediction (TCP) from observational data supports the answer of when and how to assign multiple sequential treatments, yielding importance in various applications. Despite the progress achieved by recent advances, e.g., LSTM or Transformer based causal approaches, their capability of capturing interactions in long sequences remains to be improved in both prediction performance and running efficiency. In parallel with the development of TCP, the success of the state-space models (SSMs) has achieved remarkable progress toward long-sequence modeling with saved running time. Consequently, studying how Mamba simultaneously benefits the effectiveness and efficiency of TCP becomes a compelling research direction. In this paper, we propose to exploit advantages of the SSMs to tackle the TCP task, by introducing a counterfactual Mamba model with Covariate-based Decorrelation towards Selective Parameters (Mamba-CDSP). Motivated by the over-balancing problem in TCP of the direct covariate balancing methods, we propose to de-correlate between the current treatment and the representation of historical covariates, treatments, and outcomes, which can mitigate the confounding bias while preserve more covariate information. In addition, we show that the overall de-correlation in TCP is equivalent to regularizing the selective parameters of Mamba over each time step, which leads our approach to be effective and lightweight. We conducted extensive experiments on both synthetic and real-world datasets, demonstrating that Mamba-CDSP not only outperforms baselines by a large margin, but also exhibits prominent running efficiency. Haotian Wang 0001, Haoxuan Li 0001, Hao Zou 0001, Haoang Chi, Long Lan, Wanrong Huang, Wenjing Yang 0002 |
ICLR | 6 |
| 2025 | AGFT-Tracker: Adaptive Game-Based PEFT for Object Tracking with PLMsabstractThe rise of pre-trained large models (PLMs) has sparked interest in vision tasks like object tracking. However, as PLMs scale, fully fine-tuning all parameters becomes impractical, highlighting the need for parameter-efficient fine-tuning (PEFT). While adapter tuning, which adds tunable parameters to Multi-Head Attention (MHA) or Feed-Forward Networks (FFN), is common, critical parameters like Layer Normalization (LN), vital for stability and convergence, are often overlooked. Furthermore, traditional fine-tuning strategies fail to differentiate module importance, limiting performance improvements. To solve these issues, we propose a new PEFT method for unlocking large model potential in object tracking: Adaptive Game-Based Fine-tuning Tracker (AGFT-Tracker). AGFT-Tracker combines adapter tuning with direct LN fine-tuning and adaptively allocates parameter budgets based on tracking attention losses. Important sensitive modules use higher-rank LoRA and frozen LN, while stable modules undergo lower-rank LoRA and LN adjustments. This approach improves effectiveness and efficiency, achieving state-of-the-art results on challenging benchmarks. Mingyu Cao, Xihuai He, Xueqiong Li, Kedi Zhang, Yuhua Tang, Wanrong Huang, Huibin Tan |
ICME | 6 |
| 2025 | Adaptive Distribution-Aware Modeling for Transformer TrackingabstractAdapting to changes in data distribution is a major challenge in visual object tracking. In Transformer-based tracking, Layer Normalization (LN) is often applied uniformly to both template and search features, limiting feature diversity. Additionally, models tend to converge to trivial solutions, and tracking samples are sensitive to distribution shifts, affecting robustness. To address these issues, we propose the Adaptive Distribution-Aware Transformer Tracker (ADAT), incorporating three key components: the Target-Aware Module (TAM), the Region-Aware Module (RAM), and the Self-Feedback-Aware Module (SFAM). TAM normalizes template and search features separately, preserving flexibility and enhancing target learning. RAM refines target perception by distinguishing between near and far target regions. SFAM filters out noisy samples and fine-tunes normalization parameters through self-feedback. While TAM and RAM regulate feature-level distribution, SFAM adjusts at the sample level. Extensive experiments show that ADAT outperforms existing methods, achieving superior performance on challenging benchmarks. Mingyu Cao, Huibin Tan, Xueqiong Li, Wanrong Huang, Kedi Zhang, Yuhua Tang, Shaowu Yang |
ICME | 4 |
| 2025 | Multi-Resolution Infrared-Visible Image Fusion using Multi-Scale Residual QuantizationabstractInfrared-visible image fusion (IVF) is an essential task in multimodal image processing that integrates infrared and visible modalities to enhance the overall image information content. However, existing methods often suffer from limited precision and efficiency. Furthermore, they fail to address practical requirements such as multi-resolution fusion and mutual translation. In this paper, we propose the Multi-Scale Residual Quantized Infrared-Visible Image Fusion (M-RQIVF) framework to efficiently generate high-quality fusion images. M-RQIVF trains multi-scale residual quantized infrared and visible autoencoders that convert images into multi-scale discrete token maps. This approach approximates the residuals from the features on a scale-by-scale basis, allowing for coarse-to-fine fused image generation that aligns well with human visual perception. Furthermore, by leveraging these discrete token maps, we train Visual Auto-Regressive (VAR) transformers using next-scale prediction. The VAR transformer ensures that features of corresponding sizes can be generated, even when the input infrared and visible images have different resolutions, facilitating fine-grained fusion. Additionally, the autoregressive structure enables image translation to be treated as a conditional generation task, thereby enabling mutual translation between infrared and visible images. Extensive experiments demonstrate that M-RQIVF outperforms the SOTAs while maintaining a much faster inference speed. Huibin Tan, Wanrong Huang, Yuhua Tang, Xueqiong Li |
ICME | 4 |
| 2025 | Transformer-Based Spatial-Temporal Counterfactual Outcomes EstimationabstractThe real world naturally has dimensions of time and space. Therefore, estimating the counterfactual outcomes with spatial-temporal attributes is a crucial problem. However, previous methods are based on classical statistical models, which still have limitations in performance and generalization. This paper proposes a novel framework for estimating counterfactual outcomes with spatial-temporal attributes using the Transformer, exhibiting stronger estimation ability. Under mild assumptions, the proposed estimator within this framework is consistent and asymptotically normal. To validate the effectiveness of our approach, we conduct simulation experiments and real data experiments. Simulation experiments show that our estimator has a stronger estimation capability than baseline methods. Real data experiments provide a valuable conclusion to the causal effect of conflicts on forest loss in Colombia. The source code is available at this [URL](https://github.com/lihe-maxsize/DeppSTCI_Release_Version-master). Haoang Chi, Wanrong Huang, Wenjing Yang 0002 |
ICML | 4 |
| 2025 | Breaking the Gradient Barrier: Unveiling Large Language Models for Strategic ClassificationabstractStrategic classification (SC) explores how individuals or entities modify their features strategically to achieve favorable classification outcomes. However, existing SC methods, which are largely based on linear models or shallow neural networks, face significant limitations in terms of scalability and capacity when applied to real-world datasets with significantly increasing scale, especially in financial services and the internet sector.
In this paper, we investigate how to leverage large language models to design a more scalable and efficient SC framework, especially in the case of growing individuals engaged with decision-making processes. Specifically, we introduce GLIM, a gradient-free SC method grounded in in-context learning.
During the feed-forward process of self-attention, GLIM implicitly simulates the typical bi-level optimization process of SC, including both the feature manipulation and decision rule optimization.
Without fine-tuning the LLMs, our proposed GLIM enjoys the advantage of cost-effective adaptation in dynamic strategic environments. Theoretically, we prove GLIM can support pre-trained LLMs to adapt to a broad range of strategic manipulations. We validate our approach through experiments with a collection of pre-trained LLMs on real-world and synthetic datasets in financial and internet domains, demonstrating that our GLIM exhibits both robustness and efficiency, and offering an effective solution for large-scale SC tasks. Xinpeng Lv, Yunxin Mao, Haoxuan Li 0001, Ke Liang 0006, Jinxuan Yang, Wanrong Huang, Haoang Chi, Long Lan, Yuanlong Chen, Wenjing Yang 0002, Haotian Wang 0001 |
NeurIPS | 6 |
| 2025 | Self-supervised re-identification for online joint multi-object trackingabstractRecently, the bottleneck of multi-object tracking is shifting from detection performance to association performance. However, research on association algorithms requires a large number of identity labels, which are more expensive than detection labels. To circumvent the need for identity labels, we propose a Self-supervised Re-identification module for online joint Multi-Object Tracking (SR-MOT). Specifically, we design an appearance discriminator to judge identities based solely on detection hypotheses and then associate the same identity with the final trajectory. To train the discriminator without using identity labels, we construct negative pairs by the detections that appear in the same video frame, as they definitely belong to different identities. Positive pairs are naturally constructed through several useful data augmentation strategies at the box level. In addition, our proposed method balances conflicting detection and re-ID tasks by using different output features and dynamically adjusts detection and re-ID loss weights based on the information content of the loss distribution to promote balance between the two tasks from the feature level and optimization methods. In our evaluation on the MOT Challenge benchmark, we show that our SR-MOT performs comparably to supervised methods and is significantly superior to other unsupervised methods. Our proposed method provides a practical solution for multi-object tracking without the need for identity labels, making it more accessible for real-world applications. Shuman Li, Longqi Yang 0002, Huibin Tan, Binglin Wang, Wanrong Huang, Hengzhu Liu, Wenjing Yang 0002, Long Lan |
Knowl. Inf. Syst. | 5 |
| 2025 | BIRDNN: Behavior-Imitation Based Repair for Deep Neural Networks
Taoran Wu, Changyuan Zhao, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002, Ji Wang 0001, Wanrong Huang |
Neural Networks | 8 |
| 2024 | Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer TrackingabstractRegarded as a template-matching task for a long time, visual object tracking has witnessed significant progress in space-wise exploration. However, since tracking is performed on videos with substantial time-wise information, it is important to simultaneously mine the temporal contexts which have not yet been deeply explored. Previous supervised works mostly consider template reform as the breakthrough point, but they are often limited by additional computational burdens or the quality of chosen templates. To address this issue, we propose a Space-Time Consistent Transformer Tracker (STCFormer), which uses a sequential fusion framework with multi-granularity consistency constraints to learn spatiotemporal context information. We design a sequential fusion framework that recombines template and search images based on tracking results from chronological frames, fusing updated tracking states in training. To further overcome the over-reliance on the fixed template without increasing computational complexity, we design three space-time consistent constraints: Label Consistency Loss (LCL) for label-level consistency, Attention Consistency Loss (ACL) for patch-level ROI consistency, and Semantic Consistency Loss (SCL) for feature-level semantic consistency. Specifically, in ACL and SCL, the label information is used to constrain the attention and feature consistency of the target and the background, respectively, to avoid mutual interference. Extensive experiments have shown that our STCFormer outperforms many of the best-performing trackers on several popular benchmarks. Wenjing Yang 0002, Wanrong Huang, Xianchen Zhou, Mingyu Cao, Huibin Tan |
AAAI | 3 |
| 2024 | Learning to Learn Better Visual PromptsabstractPrompt tuning provides a low-cost way of adapting vision-language models (VLMs) for various downstream vision tasks without requiring updating the huge pre-trained parameters. Dispensing with the conventional manual crafting of prompts, the recent prompt tuning method of Context Optimization (CoOp) introduces adaptable vectors as text prompts. Nevertheless, several previous works point out that the CoOp-based approaches are easy to overfit to the base classes and hard to generalize to novel classes. In this paper, we reckon that the prompt tuning works well only in the base classes because of the limited capacity of the adaptable vectors. The scale of the pre-trained model is hundreds times the scale of the adaptable vector, thus the learned vector has a very limited ability to absorb the knowledge of novel classes. To minimize this excessive overfitting of textual knowledge on the base class, we view prompt tuning as learning to learn (LoL) and learn the prompt in the way of meta-learning, the training manner of dividing the base classes into many different subclasses could fully exert the limited capacity of prompt tuning and thus transfer it power to recognize the novel classes. To be specific, we initially perform fine-tuning on the base class based on the CoOp method for pre-trained CLIP. Subsequently, predicated on the fine-tuned CLIP model, we carry out further fine-tuning in an N-way K-shot manner from the perspective of meta-learning on the base classes. We finally apply the learned textual vector and VLM for unseen classes.Extensive experiments on benchmark datasets validate the efficacy of our meta-learning-informed prompt tuning, affirming its role as a robust optimization strategy for VLMs. Fengxiang Wang 0004, Wanrong Huang, Shaowu Yang, Long Lan |
AAAI | 2 |
| 2024 | Modality Re-Balance for Visual Question Answering: A Causal FrameworkabstractVisual Question Answering (VQA) models often prioritize language cues over visual knowledge, leading to the "language prior" phenomenon. To address this, researchers have proposed methods to balance language and image information during training and inference. However, these approaches often struggle to capture important linguistic components due to the excessive exclusion of language information. Inspired by causal inference, we introduce a novel approach called the SyMmetrically Balanced Causal framework (SMBC) that rebalances visual and textual information in VQA tasks. This framework allows for an equal contribution of knowledge from both modalities to inference results. Experimental evaluation shows that SMBC: 1) applies to prevalent VQA models, including those with data augmentation, and 2) consistently improves performance on established benchmarks. Xinpeng Lv, Wanrong Huang, Haotian Wang 0001, Ruochun Jin, Xueqiong Li, Shuman Li, Yongquan Feng, Yuhua Tang |
ICASSP | 2 |
| 2024 | Boosting Meaningful Dependency Mining with Clustering and Covariance AnalysisabstractFunctional dependencies (FDs) form a valuable ingredient for various data management tasks. However, existing methods can hardly discover practical and interpretable FDs, especially in large noisy real-life datasets. This paper studies the problem of discovering meaningful functional dependencies (FDms) that utilize support and error parameters to capture interesting dependencies in such datasets and proposes an efficient discovery algorithm called FDMε. In order to scale with large datasets, FDM ε employs an efficient sampling method with accuracy guarantees to capture the differences between tuple pairs and to quantify the connection between support/error of dependencies on samples and those on the entire dataset. Moreover, it adopts a clustering-based correlated attributes extraction to divide the exponentially large search space into multiple small sub-spaces and proposes an easy-first traversal strategy with covariance-based guidance that quickly detects candidate dependencies and validates them. Additionally, we prove a covariance lower bound as an additional pruning criterion to reduce the search space. Extensive experiments on real-life and synthetic datasets demonstrate that FDM ε is 14 times faster than existing discovery algorithms on average, up to 31 times, and scales to larger datasets with the least memory cost. Ruochun Jin, Wanrong Huang, Yuhua Tang |
ICDE | 3 |
| 2024 | Out-of-Distribution Generalization With Causal Feature SeparationabstractDriven by empirical risk minimization, machine learning algorithm tends to exploit subtle statistical correlations existing in the training environment for prediction, while the spurious correlations are unstable across environments, leading to poor generalization performance. Accordingly, the problem of the Out-of-distribution (OOD) generalization aims to exploit an invariant/stable relationship between features and outcomes that generalizes well on all possible environments. To address the spurious correlation induced by the selection bias, in this article, we propose a novel Clique-based Causal Feature Separation (CCFS) algorithm by explicitly incorporating the causal structure to identify causal features of outcome for OOD generalization. Specifically, the proposed CCFS algorithm identifies the largest clique in the learned causal skeleton. Theoretically, we guarantee that either the largest clique or the rest of the causal skeleton is exactly the set of all causal features of the outcome. Finally, we separate the causal features from the non-causal ones with a sample-reweighting decorrelator for OOD prediction. Extensive experiments validate the effectiveness of the proposed CCFS method on both causal feature identification and OOD generalization tasks. Haotian Wang 0001, Kun Kuang 0001, Long Lan, Zige Wang, Wanrong Huang, Fei Wu 0001, Wenjing Yang 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | OL-CBBA: An Online Task Allocation Algorithm under Weak Communication ConditionsabstractMultiple unmanned aerial vehicles (UAVs) and multiple tasks allocation problem is difficult to solve. Existing task allocation algorithms assume that the UAVs’ position is static, and cannot assign tasks with the changing UAVs’ position simultaneously during task execution. Those algorithms reduce the efficiency of task allocation. In this paper, we propose an online consensus-based bundle algorithm (OL-CBBA) under weak communication for dynamic task allocation. We consider the situation that the location information of UAVs will constantly change during the dynamic execution of tasks. The algorithm first improves the static CBBA to an online algorithm by updating the task marginal score in task path. Moreover, we specify a flag for the convergence of individual tasks, allowing UAVs to start executing tasks earlier. Extensive comparative experiments prove the highly consistent efficiency of OL-CBBA under weak communication conditions. Specifically, the proposed OL-CBBA attains up to 22% improvement compared with CBBA. Wanrong Huang, Zhongxuan Cai |
ICPADS | 2 |
| 2023 | Treatment Effect Estimation with Adjustment Feature SelectionabstractIn causal inference, it is common to select a subset of observed covariates, named the adjustment features, to be adjusted for estimating the treatment effect. For real-world applications, the abundant covariates are usually observed, which contain extra variables partially correlating to the treatment (treatment-only variables, e.g., instrumental variables) or the outcome (outcome-only variables, e.g., precision variables) besides the confounders (variables that affect both the treatment and outcome). In principle, unbiased treatment effect estimation is achieved once the adjustment features contain all the confounders. However, the performance of empirical estimations varies a lot with different extra variables. To solve this issue, variable separation/selection for treatment effect estimation has received growing attention when the extra variables contain instrumental variables and precision variables. Haotian Wang 0001, Kun Kuang 0001, Haoang Chi, Longqi Yang 0002, Mingyang Geng, Wanrong Huang, Wenjing Yang 0002 |
KDD | 6 |
| 2023 | Null-text Guidance in Diffusion Models is Secretly a Cartoon-style CreatorabstractClassifier-free guidance is an effective sampling technique in diffusion models that has been widely adopted. The main idea is to extrapolate the model in the direction of text guidance and away from null-text guidance. In this paper, we demonstrate that null-text guidance in diffusion models is secretly a cartoon-style creator, i.e., the generated images can be efficiently transformed into cartoons by simply perturbing the null-text guidance. Specifically, we proposed two disturbance methods, i.e., Rollback disturbance (Back-D) and Image disturbance (Image-D), to construct misalignment between the noisy images used for predicting null-text guidance and text guidance (subsequently referred to as null-text noisy image and text noisy imageb respectively) in the sampling process. Back-D achieves cartoonization by altering the noisb level of the null-text noisy image via replacing xt with xl + Δ t. Image-D, alternatively, produces high-fidelity, diverse cartoons by defining xt as a clean input image, which further improves the incorporation of finer image details. Through comprehensive experiments, we delved into the principle of noise disturbing for null-text and uncovered that the efficacy of disturbance depends on the correlation between the null-text noisy image and the source image. Moreover, the proposed methods, which can generate cartoon images and cartoonize specific ones, are training-free and easily integrated as a plug-and-play component in any classifier-free guided diffusion model. The project page is available at https://nulltextforcartoon.github.io/. Heliang Zheng, Long Lan, Wanrong Huang, Wenjing Yang 0002 |
ACM Multimedia | 5 |
| 2023 | Meta attention for Off-Policy Actor-Critic
Jiateng Huang, Wanrong Huang, Long Lan |
Neural Networks | 2 |
| 2023 | Recent Advances for Quantum Neural Networks in Generative LearningabstractQuantum computers are next-generation devices that hold promise to perform calculations beyond the reach of classical computers. A leading method towards achieving this goal is through quantum machine learning, especially quantum generative learning. Due to the intrinsic probabilistic nature of quantum mechanics, it is reasonable to postulate that quantum generative learning models (QGLMs) may surpass their classical counterparts. As such, QGLMs are receiving growing attention from the quantum physics and computer science communities, where various QGLMs that can be efficiently implemented on near-term quantum machines with potential computational advantages are proposed. In this paper, we review the current progress of QGLMs from the perspective of machine learning. Particularly, we interpret these QGLMs, covering quantum circuit Born machines, quantum generative adversarial networks, quantum Boltzmann machines, and quantum variational autoencoders, as the quantum extension of classical generative learning models. In this context, we explore their intrinsic relations and their fundamental differences. We further summarize the potential applications of QGLMs in both conventional machine learning tasks and quantum physics. Last, we discuss the challenges and further research directions for QGLMs. Jinkai Tian, Shanshan Zhao 0001, Qing Liu 0027, Kaining Zhang, Wanrong Huang, Xingyao Wu, Min-Hsiu Hsieh, Tongliang Liu, Wenjing Yang 0002, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2021 | BT Expansion: a Sound and Complete Algorithm for Behavior Planning of Intelligent Robots with Behavior TreesabstractBehavior Trees (BTs) have attracted much attention in the robotics field in recent years, which generalize existing control architectures and bring unique advantages for building robot systems. Automated synthesis of BTs can reduce human workload and build behavior models for complex tasks beyond the ability of human design, but theoretical studies are almost missing in existing methods because it is difficult to conduct formal analysis with the classic BT representations. As a result, they may fail in tasks that are actually solvable. This paper proposes BT expansion, an automated planning approach to building intelligent robot behaviors with BTs, and proves the soundness and completeness through the state-space formulation of BTs. The advantages of blended reactive planning and acting are formally discussed through the region of attraction of BTs, by which robots with BT expansion are robust to any resolvable external disturbances. Experiments with a mobile manipulator and test sets are simulated to validate the effectiveness and efficiency, where the proposed algorithm surpasses the baseline by virtue of its soundness and completeness. To the best of our knowledge, it is the first time to leverage the state-space formulation to synthesize BTs with a complete theoretical basis. Zhongxuan Cai, Minglong Li, Wanrong Huang, Wenjing Yang 0002 |
AAAI | 3 |
| 2021 | Model Compression for a Plasticity Neural Network in a Maze Exploration Scenario
Baolun Yu, Wanrong Huang, Long Lan, Yuhua Tang |
ICONIP (5) | 2 |
| 2020 | One-shot video-based person re-identification with variance subsampling algorithmabstractAbstract Previous works propose the distance‐based sampling for unlabeled datapoints to address the few‐shot person re‐identification task, however, many selected samples may be assigned with wrong labels due to poor feature quality in these works, which negatively affects the learning procedure. In this article, we propose a novel sampling strategy to improve the quality of assigned pseudo‐labels, thus promoting the final performance. To illustrate, we first propose the concept of variance confidence to measure the credibility of pseudo‐labels, then we apply a novel variance subsampling algorithm to improve the accuracy of the selected sample labels. Our method combines distance confidence and variance confidence as a two‐round sampling criterion. Meanwhile, a variation decay strategy is used in our sampling process in combination with the actual distribution of features. We evaluate our approach on two publicly available datasets, MARS and DukeMTMC‐VideoReID, and achieve state‐of‐the‐art one‐shot performance. Wenjing Yang 0002, Wanrong Huang, Qiong Yang |
Comput. Animat. Virtual Worlds | 4 |
| 2018 | Multi-feature Fusion for Deep Reinforcement Learning: Sequential Control of Mobile Robots
Haotian Wang 0001, Wenjing Yang 0002, Wanrong Huang, Yuhua Tang |
ICONIP (7) | 3 |
| 2018 | Distributed coordination with connectivity maintenance for nonholonomic robotsabstractAbstract Multirobot systems have been studied extensively in the recent years. Maintaining connectivity has significant impacts on the stability and convergence of the multirobot systems. In this work, we design a three‐layer framework for multirobot coordination. Furthermore, a novel distributed algorithm is proposed to achieve the navigation objective while satisfying connectivity maintenance and collision avoidance constraints. The algorithm is a hybrid of an rapidly exploring random tree‐based planner and an extended distributed navigation function‐based controller. The coordination framework and the distributed algorithm are demonstrated to be effective through a series of illustrative simulations. They outperform the current state‐of‐the‐art method in terms of efficiency and applicability. Wanrong Huang, Xiaodong Yi 0002, Xuejun Yang |
Comput. Animat. Virtual Worlds | 1 |
| 2018 | Distributed coordination with connectivity maintenance for nonholonomic robotsabstractSubsequent to publication, the affiliation and citation of the article by Huang et al.1 have been modified. The correct order is presented above. Wanrong Huang, Xiaodong Yi 0002, Xuejun Yang |
Comput. Animat. Virtual Worlds | 1 |
| 2016 | Multi-level Occupancy Grids for Efficient Representation of 3D Indoor Environments
Wanrong Huang, Xiaodong Yi 0002, Xuejun Yang |
PRICAI | 2 |