Hao Chen 0103

dblp:175/3324-103 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0002-5982-0615ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Time Series Supplier Allocation via Deep Black-Litterman Model
abstract
As a typical problem of Spatiotemporal Resource Management, Time Series Supplier Allocation (TSSA) poses a complex NP-hard challenge, aimed at refining future order dispatching strategies to satisfy the trade-off between demands and maximum supply. The Black-Litterman (BL) model, which comes from financial portfolio management, offers a new perspective for the TSSA by balancing expected returns against insufficient supply risks. However, the BL model is not only constrained by manually constructed perspective matrices and spatio-temporal market dynamics but also restricted by the absence of supervisory signals and unreliable supplier data. To solve these limitations, we introduce the pioneering Deep Black-Litterman Model for TSSA, which innovatively adapts the BL model from financial domain to supply chain context. Specifically, DBLM leverages Spatio-Temporal Graph Neural Networks (STGNNs) to capture spatio-temporal dependencies for automatically generating future perspective matrices. Moreover, a novel Spearman rank correlation is designed as our DBLM supervise signal to navigate complex risks and interactions of the supplier. Finally, DBLM further uses a masking mechanism to counteract the bias of unreliable data, thus improving precision and reliability. Extensive experiments on two datasets demonstrate significant improvements of DBLM on TSSA.
Xinke Jiang, Wentao Zhang 0008, Yuchen Fang 0001, Hao Chen 0103, Dingyi Zhuang, Jiayuan Luo
AAAI5
2025 TC-RAG: Turing-Complete RAG's Case study on Medical LLM Systems
abstract
Xinke Jiang, Yue Fang, Rihong Qiu, Haoyu Zhang, Yongxin Xu, Hao Chen, Wentao Zhang, Ruizhe Zhang, Yuchen Fang, Xinyu Ma, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinke Jiang, Rihong Qiu, Yongxin Xu, Hao Chen 0103, Wentao Zhang 0008, Ruizhe Zhang 0013, Yuchen Fang 0001, Junfeng Zhao 0001, Yasha Wang
ACL (1)6
2025 Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
abstract
Pre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As the practical applications of VLMs become increasingly widespread, their potential safety and robustness issues raise concerns that adversaries may evade the system and cause these models to generate toxic content through malicious attacks. Therefore, evaluating the robustness of open-source VLMs against adversarial attacks has garnered growing attention, with transfer-based attacks as a representative black-box attacking strategy. However, most existing transfer-based attacks neglect the importance of the semantic correlations between vision and text modalities, leading to sub-optimal adversarial example generation and attack performance. To address this issue, we present Chain of Attack (CoA)1, which iteratively enhances the generation of adversarial examples based on the multi-modal semantic update using a series of intermediate attacking steps, achieving superior adversarial transferability and efficiency. A unified attack success rate computing method is further proposed for automatic evasion evaluation. Extensive experiments conducted under the most realistic and high-stakes scenario, demonstrate that our attacking strategy is able to effectively mislead models to generate targeted responses using only black-box attacks without any knowledge of the victim models. The comprehensive robustness evaluation in our paper provides insight into the vulnerabilities of VLMs and offers a reference for the safety considerations of future model developments.
Yequan Bie, Jianda Mao, Yangqiu Song, Yang Wang 0020, Hao Chen 0103, Kani Chen
CVPR6
2025 GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
abstract
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS.
Tianhao Chen, Xin Xu 0001, Zijing Liu, Xinyuan Song 0002, Ajay Jaiswal, Jishan Hu, Yang Wang 0020, Hao Chen 0103, Shizhe Diao, Shiwei Liu 0003, Lu Yin 0006, Can Yang 0002
NeurIPS10
2025 SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
abstract
Code-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phenomenon poses challenges for Automatic Speech Recognition (ASR) systems, which are typically designed for a single language and struggle to handle multilingual inputs. The growing global demand for multilingual applications, including Code-Switching ASR (CSASR), Text-to-Speech (TTS), and Cross-Lingual Information Retrieval (CLIR), highlights the inadequacy of existing monolingual datasets. Although some code-switching datasets exist, most are limited to bilingual mixing within homogeneous ethnic groups, leaving a critical need for a large-scale, diverse benchmark akin to ImageNet in computer vision. To bridge this gap, we introduce \textbf{LinguaMaster}, a multi-agent collaboration framework specifically designed for efficient and scalable multilingual data synthesis. Leveraging this framework, we curate \textbf{SwitchLingua}, the first large-scale multilingual and multi-ethnic code-switching dataset, including: (1) 420K CS textual samples across 12 languages, and (2) over 80 hours of audio recordings from 174 speakers representing 18 countries/regions and 63 racial/ethnic backgrounds, based on the textual data. This dataset captures rich linguistic and cultural diversity, offering a foundational resource for advancing multilingual and multicultural research. Furthermore, to address the issue that existing ASR evaluation metrics lack sensitivity to code-switching scenarios, we propose the \textbf{Semantic-Aware Error Rate (SAER)}, a novel evaluation metric that incorporates semantic information, providing a more accurate and context-aware assessment of system performance. Benchmark experiments on SwitchLingua with state-of-the-art ASR models reveal substantial performance gaps, underscoring the dataset’s utility as a rigorous benchmark for CS capability evaluation. In addition, SwitchLingua aims to encourage further research to promote cultural inclusivity and linguistic diversity in speech technology, fostering equitable progress in the ASR field. LinguaMaster (Code): github.com/Shelton1013/SwitchLingua, SwitchLingua (Data): https://huggingface.co/datasets/Shelton1013/SwitchLinguatext, https://huggingface.co/datasets/Shelton1013/SwitchLinguaaudio
Xingyuan Liu, Yequan Bie, Tsz Wai Chan, Yangqiu Song, Yang Wang 0020, Hao Chen 0103, Kani Chen
NeurIPS7
2025 Exploration via Embracing Diversity in Reinforcement Learning for Sparse-Reward Procedurally-Generated Tasks
abstract
A key challenge in reinforcement learning is how to guide agents to efficiently explore sparse reward environments. In order to overcome this challenge, the state-of-the-art methods introduce additional intrinsic rewards based on state-related information, such as the novelty of states. Unfortunately, these methods frequently fail in procedurally-generated tasks, where a different environment is generated in each episode so that the agent is not likely to visit the same state more than once. Recently, some exploration methods designed specifically for procedurally-generated tasks have been proposed. However, they still only consider state-related information, which leads to relatively inefficient exploration. In this work, we propose a novel exploration method, which utilizes cross-episode policy-related information and intraepisode state-related information to jointly encourage exploration in procedurally-generated tasks. In term of policy-related information, we first use an imitator-based unbalanced policy diversity to measure the difference between the agent’s current policy and the agent’s previous policies, and then encourage the agent to maximize this difference. In term of state-related information, we encourage the agent to maximize the state diversity within an episode, thereby visiting as many different states as possible in an episode. We show that our method significantly improves sample efficiency over state-of-the-art methods on three challenging benchmarks, including MiniGrid, MiniWorld, and the sparse-reward version of Procgen.
Pei Xu 0003, Hao Chen 0103, Wenjie Yang 0005, Kaiqi Huang
IEEE Trans. Syst. Man Cybern. Syst.2
2024 GATE: Guided Contrastive State Space for Multi-agent Reinforcement Learning
Hao Chen 0103, Bin Zhang 0052
ICONIP (4)1
2024 SGCD: Subgroup Contribution Decomposition for Multi-Agent Reinforcement Learning
abstract
Cooperative multi-agent reinforcement learning (MARL) tasks rely on the efficient coordination among agents, working collectively as a team to address diverse challenges. However, considering the team as a cohesive entity introduces a flat structure to cooperation. In contrast, grouping serves as a method to tackle the issue by decomposing the team, thereby providing a more compact representation of the team’s structure. While grouping has been proven effective, numerous grouping methods are limited to specific composition structures and struggle to introduce diverse group patterns into the framework. In this paper, we propose SGCD, a subgroup contribution decomposition method that incorporates the idea of subgroups and inner subgroups, leveraging the Shapley Value to distribute contributions. This approach facilitates the decomposition of contributions from subgroups to the collective onto individual agents, enabling the high-level network to maintain consistency across various grouping patterns, thereby fostering continued cooperation among agents. Notably, our decomposition method is not confined to a specific team decomposition, making it adaptable to different grouping structures. The effectiveness of SGCD is demonstrated through experiments conducted in the Google Research Football (GRF) and StarCraft Multi-Agent Challenge (SMAC) environments.
Hao Chen 0103, Bin Zhang 0052
IJCNN1
2023 Consensus Learning for Cooperative Multi-Agent Reinforcement Learning
abstract
Almost all multi-agent reinforcement learning algorithms without communication follow the principle of centralized training with decentralized execution. During the centralized training, agents can be guided by the same signals, such as the global state. However, agents lack the shared signal and choose actions given local observations during execution. Inspired by viewpoint invariance and contrastive learning, we propose consensus learning for cooperative multi-agent reinforcement learning in this study. Although based on local observations, different agents can infer the same consensus in discrete spaces without communication. We feed the inferred one-hot consensus to the network of agents as an explicit input in a decentralized way, thereby fostering their cooperative spirit. With minor model modifications, our suggested framework can be extended to a variety of multi-agent reinforcement learning algorithms. Moreover, we carry out these variants on some fully cooperative tasks and get convincing results.
Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Zeren Zhang, Guangchong Zhou, Hao Chen 0103
AAAI6
2023 Uncertainty Quantification via Spatial-Temporal Tweedie Model for Zero-inflated and Long-tail Travel Demand Prediction
abstract
Understanding Origin-Destination (O-D) travel demand is crucial for transportation management. However, traditional spatial-temporal deep learning models grapple with addressing the sparse and long-tail characteristics in high-resolution O-D matrices and quantifying prediction uncertainty. This dilemma arises from the numerous zeros and over-dispersed demand patterns within these matrices, which challenge the Gaussian assumption inherent to deterministic deep learning models. To address these challenges, we propose a novel approach: the Spatial-Temporal Tweedie Graph Neural Network (STTD). The STTD introduces the Tweedie distribution as a compelling alternative to the traditional 'zero-inflated' model and leverages spatial and temporal embeddings to parameterize travel demand distributions. Our evaluations using real-world datasets highlight STTD's superiority in providing accurate predictions and precise confidence intervals, particularly in high-resolution scenarios. GitHub code is available online(https://github.com/STTDAnonymous/STTD).
Xinke Jiang, Dingyi Zhuang, Hao Chen 0103, Jiayuan Luo
CIKM4
2023 Explicitly Learning Policy Under Partial Observability in Multiagent Reinforcement Learning
abstract
We explore explicit solutions for multiagent reinforcement learning (MARL) under the constraint of partial observability. With a general framework of centralized training with decentralized execution (CTDE), existing methods implicitly alleviate partial observability by introducing global information during centralized training. However, such implicit solution cannot well address partial observability and shows low sample efficiency in many MARL problems. In this paper, we focus on the influence of partial observability on the policy of agents, and formally derive an ideal form of policy that maximizes MARL objective under partial observability. Furthermore, we develop a new method named Explicitly Learning Policy (ELP), which adopts a novel teacher-student structure and utilizes knowledge distillation to explicitly learn individual policy under partial observability for each agent. Compared to prior methods, ELP presents a more general and interpretable training process, and the procedure of ELP can be easily extended to existing methods for performance boost. Our empirical experiments on StarCraft II micromanagement benchmark show that ELP significantly outperforms prevailing state-of-the-art baselines, which demonstrates the advantage of ELP in addressing partial observability and improving sample efficiency.
Guangkai Yang, Hao Chen 0103, Junge Zhang
IJCNN3
2023 Underexplored Subspace Mining for Sparse-Reward Cooperative Multi-Agent Reinforcement Learning
abstract
Learning cooperation in sparse-reward multi-agent reinforcement learning is challenging, since agents need to explore in the large joint-state space with sparse feedback. However, in cooperative games, the cooperative target is often related to partial attributes, hence there is no need to treat the whole state space equally. Therefore, we propose Underexplored Subspace Mining (USM), a novel type of intrinsic reward that encourages agents to selectively explore partial attributes instead of wasting time on the whole state space to accelerate learning. Specially, considering that the target-related attributes are varying in different games and hard to predefine, we choose to focus on the underexplored subspace as an alternative, which is an automatic aggregation of the underexplored bottom-level dimensions without any human design or learning parameters. We evaluate our method in cooperative games with discrete and continuous state space separately. Results demonstrate that USM consistently outperforms existing state-of-the-art methods, and becomes the only method that has succeeded in sparse-reward games evaluated with larger state space or more complicated cooperation dynamics.
Yang Yu 0056, Qiyue Yin, Junge Zhang, Hao Chen 0103, Kaiqi Huang
IJCNN4
2022 Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation
abstract
As RGB-D sensors become more affordable, using RGB- D images to obtain high-accuracy 6D pose estimation results becomes a better option. State-of-the-art approaches typically use different backbones to extract features for RGB and depth images. They use a 2D CNN for RGB images and a perpixel point cloud network for depth data, as well as a fusion network for feature fusion. We find that the essential reason for using two independent backbones is the “projection breakdown” problem. In the depth image plane, the projected 3D structure of the physical world is preserved by the 1D depth value and its built-in 2D pixel coordinate (UV). Any spatial transformation that modifies UV, such as resize, flip, crop, or pooling operations in the CNN pipeline, breaks the binding between the pixel value and UV coordinate. As a consequence, the 3D structure is no longer preserved by a modified depth image or feature. To address this issue, we propose a simple yet effective method denoted as Uni6D that explicitly takes the extra UV data along with RGB-D images as input. Our method has a Unified CNN framework for 6D pose estimation with a single CNN backbone. In particular, the architecture of our method is based on Mask R-CNN with two extra heads, one named RT head for directly predicting 6D pose and the other named abc head for guiding the network to map the visible points to their coordinates in the 3D model as an auxiliary module. This end-to-end approach balances simplicity and accuracy, achieving comparable accuracy with state of the arts and 7.2x faster inference speed on the YCB-Video dataset.
Xiaoke Jiang, Donghai Li, Hao Chen 0103, Rui Zhao 0001
CVPR3
2022 Layer-Wisely Supervised Learning For One-Shot Neural Architecture Search
abstract
Neural architecture search aims to automatically discover both efficient and effective neural architectures. Recently, one-shot neural architecture search (one-shot NAS) has drawn great attention due to its high efficiency and competitive performance. One of the most important problems in one-shot NAS is to evaluate the capabilities of architecture candidates. In particular, a pre-trained super-net is served as an evaluator. Due to the large weight-sharing space, current one-shot methods suffer from the ranking disorder issue, that is, the ranking correlation between estimated capabilities and true capabilities of candidates is incorrect. Moreover, the super-net in search is dense thus it is inefficient to train with end-to-end back-propagation. In this paper, we propose to modularize the large weight-sharing space of one-shot NAS into layers by introducing layer-wisely supervised learning. But we discover that greedy layer-wise learning that learns each layer separately with a local objective hurts super-net performance as well as ranking correlation. Instead, we learn each layer by using the gradients propagated from the objective associated with the adjacent upper layer. The simple proposal reduces the representation shift and improves the ranking correlation. In addition, it reduces 47.4% memory footprint and gets a faster convergence of super-net training compared with the strong baseline. Extensive experiments on ImageNet with both supervised and self-supervised objectives demonstrate the effectiveness of our proposal.
Zhourui Guo, Qiyue Yin, Hao Chen 0103, Kaiqi Huang
IJCNN4
2022 Multi-Agent Uncertainty Sharing for Cooperative Multi-Agent Reinforcement Learning
abstract
Cooperative multi-agent reinforcement learning has been considered promising to complete many complex cooperative tasks in the real world such as coordination of robot swarms and self-driving. To promote multi-agent cooperation, Centralized Training with Decentralized Execution emerges as a popular learning paradigm due to partial observability and communication constraints during execution and computational complexity in training. Value decomposition has been known to produce competitive performance to other methods in complex environment within this paradigm such as VDN and QMIX, which approximates the global joint Q-value function with multiple local individual Q-value functions. However, existing works often neglect the uncertainty of multiple agents resulting from the partial observability and very large action space in the multi-agent setting and can only obtain the sub-optimal policy. To alleviate the limitations above, building upon the value decomposition, we propose a novel method called multi-agent uncertainty sharing (MAUS). This method utilizes the Bayesian neural network to explicitly capture the uncertainty of all agents and combines with Thompson sampling to select actions for policy learning. Besides, we impose the uncertainty-sharing mechanism among agents to stabilize training as well as coordinate the behaviors of all the agents for multi-agent cooperation. Extensive experiments on the StarCraft Multi-Agent Challenge (SMAC) environment demonstrate that our approach achieves significant performance to exceed the prior baselines and verify the effectiveness of our method.
Hao Chen 0103, Guangkai Yang, Junge Zhang, Qiyue Yin, Kaiqi Huang
IJCNN1
2022 RACA: Relation-Aware Credit Assignment for Ad-Hoc Cooperation in Multi-Agent Deep Reinforcement Learning
abstract
In recent years, reinforcement learning has faced several challenges in the multi-agent domain, such as the credit assignment issue. Value function factorization emerges as a promising way to handle the credit assignment issue under the centralized training with decentralized execution (CTDE) paradigm. However, existing value function factorization methods cannot deal with ad-hoc cooperation, that is, adapting to new configurations of teammates at test time. Specifically, these methods do not explicitly utilize the relationship between agents and cannot adapt to different sizes of inputs. To address these limitations, we propose a novel method, called Relation-Aware Credit Assignment (RACA), which achieves zero-shot generalization in ad-hoc cooperation scenarios. RACA takes advantage of a graph-based relation encoder to encode the topological structure between agents. Furthermore, RACA utilizes an attention-based observation abstraction mechanism that can generalize to an arbitrary number of teammates with a fixed number of parameters. Experiments demonstrate that our method outperforms baseline methods on the StarCraftII micromanagement benchmark and ad-hoc cooperation scenarios.
Hao Chen 0103, Guangkai Yang, Junge Zhang, Qiyue Yin, Kaiqi Huang
IJCNN1
2022 FGA-NAS: Fast Resource-Constrained Architecture Search by Greedy-ADMM Algorithm
abstract
Differentiable architecture search has demonstrated promising results in automatically designing neural network architectures with desired properties, such as high accuracy and low FLOPs. However, it suffers from a cumbersome training process, and the injection of constraints in the search phase often relies on some hand-crafted heuristic regularizers, the design of which typically requires tremendous human effort. In this paper, to address these critical challenges, we present FGA-NAS, an efficient method for resource-constrained architecture search. First, to reduce the computational cost and improve search flexibility, we propose a novel condensed search space that merges multiple parallel-placed candidates into a single one. Second, to enable the gradient-based optimization for neural architecture search (NAS) under multiple combinatorial constraints, we decompose the constrained NAS into a few simple sub-problems without introducing any heuristics by using the ADMM algorithm [1]. Then, the constrained NAS can be resolved by alternately solving the simple sub-problems. Experimental results on ImageNet show that our method can discover efficient and accurate neural network architectures that achieve the state-of-the-art by only using 0.2 GPU days.
Junge Zhang, Qiaozhe Li, Hao Chen 0103, Kaiqi Huang
IJCNN4