Shiji Song

dblp:72/5351 · DBLP profile ↗
← Back
197ranked-venue papers
7as first author
122since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 138 · 5 first-author · 91 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 35 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 1 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 23 · 1 first-author · 14 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 A Graph-Based Reinforcement Learning Method for Flexible Job Shop Scheduling with Sequence Flexibility
Erdong Yuan, Shiji Song, Xiuxian Zhong
ICIC (14)4
2026 MED-Net: Leveraging multi-visual encoding and Manhattan-distance pooling for robust medical image segmentation
Zhaozhao Su, Shiji Song, Zhiqiang Zhu, Zhihua Fang, Bowen Wang 0029
Expert Syst. Appl.3
2026 Event-Based Security Formation Control Against Dual-Terminal False Data Injection Attacks in Autonomous Mobile Robot Swarm
abstract
This work considers the security formation control for the autonomous mobile robot swarm with uncertain nonlinearity under dual-terminal false data injection attacks on both actuators and sensors. A feasible attack compensation mechanism is developed to eliminate the influence of tampering on the transmitted signals of actuators and sensors. Reconstruction of uncompromised signals is achieved through an adaptive neural network state estimator. Considering the limited network bandwidth, an event-triggered adaptive neural network security formation control strategy is proposed by integrating the compensation mechanism and the state estimator. Sufficient conditions are derived to ensure the uniformly ultimate boundedness of the formation error system. It is rigorously proven that the proposed method can avoid the Zeno phenomenon. Finally, simulation verification demonstrates the validity of theoretical results.
Zhenyu Chang, Guangdeng Zong, Shiji Song, Xudong Zhao 0001
IEEE Internet Things J.3
2026 UltraSeP: Sequence-aware pre-training for echocardiography probe movement guidance
Haojun Jiang, Zhenguo Sun, Yulin Wang 0002, Yu Sun 0020, Meng Li 0087, Shaqi Luo, Shiji Song, Gao Huang 0001
Pattern Recognit.10
2025 Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
abstract
Test-Time adaptation (TTA) aims to improve the performance of source-domain pre-trained models on previously unseen, shifted target domains. Traditional TTA methods primarily adapt model weights based on target data streams, making model performance sensitive to the amount and order of target data. The recently proposed diffusion-driven TTA methods mitigate this by adapting model inputs instead of weights, where an unconditional diffusion model, trained on the source domain, transforms target-domain data into a synthetic domain that is expected to approximate the source domain. However, in this paper, we reveal that although the synthetic data in diffusion-driven TTA seems indistinguishable from the source data, it is unaligned with, or even markedly different from the latter for deep networks. To address this issue, we propose a Synthetic-Domain Alignment (SDA) framework. Our key insight is to fine-tune the source model with synthetic data to ensure better alignment. Specifically, we first employ a conditional diffusion model to generate labeled samples, creating a synthetic dataset. Subsequently, we use the aforementioned unconditional diffusion model to add noise to and denoise each sample before fine-tuning. This Mix of Diffusion (MoD) process mitigates the potential domain misalignment between the conditional and unconditional models. Extensive experiments across classifiers, segmenters, and multimodal large language models (MLLMs, e.g., LLaVA) demonstrate that SDA achieves superior domain alignment and consistently outperforms existing diffusion-driven TTA methods. Our code is available at https://github.com/SHI-Labs/Diffusion-Driven-Test-Time-Adaptation-Via-Synthetic-Domain-Alignment.
Junhao Zhao, Chaoqun Du, Yulin Wang 0002, Chunjiang Ge, Zanlin Ni, Shiji Song, Humphrey Shi, Gao Huang 0001
CVPR7
2025 EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
abstract
Echocardiography is crucial for cardiovascular disease detection but relies heavily on experienced sonographers. Echocardiography probe guidance systems, which provide real-time movement instructions for acquiring standard plane images, offer a promising solution for AI-assisted or fully autonomous scanning. However, developing effective machine learning models for this task remains challenging, as they must grasp heart anatomy and the intricate interplay between probe motion and visual signals. To address this, we present EchoWorld, a motion-aware world modeling framework for probe guidance that encodes anatomical knowledge and motion-induced visual dynamics, while effectively leveraging past visual-motion sequences to enhance guidance precision. EchoWorld employs a pre-training strategy inspired by world modeling principles, where the model predicts masked anatomical regions and simulates the visual outcomes of probe adjustments. Built upon this pre-trained model, we introduce a motion-aware attention mechanism in the fine-tuning stage that effectively integrates historical visual-motion data, enabling precise and adaptive probe guidance. Trained on more than one million ultrasound images from over 200 routine scans, EchoWorld effectively captures key echocar-diographic knowledge, as validated by qualitative analysis. Moreover, our method significantly reduces guidance errors compared to existing visual backbones and guidance frameworks, excelling in both single-frame and sequential evaluation protocols. Code is available at https://github.com/LeapLabTHU/EchoWorld.
Yulin Wang 0002, Haojun Jiang, Pan Liu 0004, Shiji Song, Gao Huang 0001
CVPR5
2025 CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
abstract
Humans can develop internal world models that encode common sense knowledge, telling them how the world works and predicting the consequences of their actions. This concept has emerged as a promising direction for establishing general-purpose machine-learning models in recent preliminary works, e.g., for visual representation learning. In this paper, we present CheXWorld, the first effort towards a self-supervised world model for radiographic images. Specifically, our work develops a unified framework that simultaneously models three aspects of medical knowledge essential for qualified radiologists, including 1) local anatomical structures describing the fine-grained characteristics of local tissues (e.g., architectures, shapes, and textures); 2) global anatomical layouts describing the global organization of the human body (e.g., layouts of organs and skeletons); and 3) domain variations that encourage CheXWorld to model the transitions across different appearance domains of radiographs (e.g., varying clarity, contrast, and exposure caused by collecting radiographs from different hospitals, devices, or patients). Empirically, we design tailored qualitative and quantitative analyses, revealing that CheX-World successfully captures these three dimensions of medical knowledge. Furthermore, transfer learning experiments across eight medical image classification and segmentation benchmarks showcase that CheXWorld significantly outperforms existing SSL methods and large-scale medical foundation models. Code & pre-trained models are available at https://github.com/LeapLabTHU/CheXWorld.
Yulin Wang 0002, Chenxin Tao, Pan Liu 0004, Shiji Song, Gao Huang 0001
CVPR5
2025 GridMix: Exploring Spatial Modulation for Neural Fields in PDE Modeling
abstract
Significant advancements have been achieved in PDE modeling using neural fields. Despite their effectiveness, existing methods rely on global modulation, limiting their ability to reconstruct local details. While spatial modulation with vanilla grid-based representations offers a promising alternative, it struggles with inadequate global information modeling and over-fitting to the training spatial domain. To address these challenges, we propose GridMix, a novel approach that models spatial modulation as a mixture of grid-based representations. GridMix effectively explores global structures while preserving locality for fine-grained modulation. Furthermore, we introduce spatial domain augmentation to enhance the robustness of the modulated neural fields against spatial domain variations. With all these innovations, our comprehensive approach culminates in MARBLE, a framework that significantly advancing the capabilities of neural fields in PDE modeling. The effectiveness of MARBLE is extensively validated on diverse benchmarks encompassing dynamics modeling and geometric prediction.
Shiji Song, Gao Huang 0001
ICLR2
2025 Learning to Stabilize Column Generation
abstract
Column generation is a widely adopted technique for solving linear programming problems with a large number of variables. However, standard column generation often suffers from slow convergence due to the dual solution instability. In this paper, we present a novel learning-based stabilization approach for column generation. Unlike traditional methods that address dual solution stabilization at each iteration in isolation, our method adopts a holistic perspective, leveraging its learning-based nature to explore for optimal stabilization policies that lead to faster overall convergence. We frame dual solution stabilization as a sequential decision-making problem and cast column generation as a Markov decision process. A graph convolutional neural network-based agent is employed to improve dual solution quality at each iteration. Additionally, we introduce a two-stage training scheme that combines supervised learning and reinforcement learning, ensuring stable and efficient training of the agent. Experimental evaluations on cutting stock and vertex coloring problems demonstrate that our approach outperforms several well-known stabilization methods in terms of iteration efficiency and exhibits competitive performance in terms of total runtime. Furthermore, our method shows strong generalization capabilities, performing well on significantly larger problem instances and diverse benchmarks.
Lichang Fang, Haofeng Yuan, Shiji Song, Bokui Chen
IJCNN3
2025 Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing
abstract
Huanqian Wang, Yang Yue, Rui Lu, Jingxin Shi, Andrew Zhao, Shenzhi Wang, Shiji Song, Gao Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Huanqian Wang, Rui Lu 0001, Jingxin Shi, Andrew Zhao, Shenzhi Wang, Shiji Song, Gao Huang 0001
NAACL (Long Papers)7
2025 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated notable success in enhancing the reasoning performance of large language models (LLMs), particularly in mathematics and programming tasks. It is widely believed that, similar to how traditional RL helps agents to explore and learn new strategies, RLVR enables LLMs to continuously self-improve, thus acquiring novel reasoning abilities that exceed the capacity of the corresponding base models. In this study, we take a critical look at \textit{the current state of RLVR} by systematically probing the reasoning capability boundaries of RLVR-trained LLMs across diverse model families, RL algorithms, and math/coding/visual reasoning benchmarks, using pass@\textit{k} at large \textit{k} values as the evaluation metric. While RLVR improves sampling efficiency towards the correct path, we surprisingly find that current training does \emph{not} elicit fundamentally new reasoning patterns. We observe that while RLVR-trained models outperform their base models at smaller values of $k$ (\eg, $k$=1), base models achieve higher pass@$k$ score when $k$ is large. Moreover, we observe that the reasoning capability boundary of LLMs often narrows as RLVR training progresses. Further coverage and perplexity analysis shows that the reasoning paths generated by RLVR models are already included in the base models' sampling distribution, suggesting that their reasoning abilities originate from and are \textit{bounded} by the base model. From this perspective, treating the base model as an upper bound, our quantitative analysis shows that six popular RLVR algorithms perform similarly and remain far from optimal in fully leveraging the potential of the base model. In contrast, we find that distillation can introduce new reasoning patterns from the teacher and genuinely expand the model’s reasoning capabilities. Taken together, our findings suggest that current RLVR methods have not fully realized the potential of RL to elicit genuinely novel reasoning abilities in LLMs. This underscores the need for improved RL paradigms—such as continual scaling and multi-turn agent-environment interaction—to unlock this potential.
Rui Lu 0001, Andrew Zhao, Zhaokai Wang, Shiji Song, Gao Huang 0001
NeurIPS6
2025 Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), yet its underlying mechanisms remain insufficiently understood. In this work, we undertake a pioneering exploration of RLVR through the novel perspective of token entropy patterns, comprehensively analyzing how different tokens influence reasoning performance. By examining token entropy patterns in Chain-of-Thought (CoT) reasoning, we observe that only a small fraction (approximately 20\%) of tokens exhibit high entropy, and these tokens semantically act as critical forks that steer the model toward diverse reasoning pathways. We further demonstrate that moderately increasing the entropy of these high-entropy tokens via decoding temperature adjustments leads to improved performance, quantitatively confirming their role as decision points in reasoning. We ultimately refine RLVR by restricting policy gradient updates to these forking tokens. Despite utilizing only 20\% of tokens, our approach achieves comparable performance to full-gradient updates on the Qwen3-8B base model. Moreover, it demonstrates remarkable improvements on the larger Qwen3-32B base model, boosting AIME'25 scores by 11.04 and AIME'24 scores by 7.71. In contrast, training exclusively on the 80\% lowest-entropy tokens leads to a marked decline in performance. These findings indicate that the efficacy of RLVR primarily arises from optimizing the high-entropy tokens that dictate key reasoning directions. Collectively, our results suggest promising avenues for optimizing RLVR algorithms by strategically leveraging the potential of these high-entropy minority tokens to further enhance the reasoning abilities of LLMs.
Shenzhi Wang, Chujie Zheng, Rui Lu 0001, Kai Dang, Xiong-Hui Chen, Jianxin Yang, Zhenru Zhang, Yuqiong Liu, An Yang, Andrew Zhao, Shiji Song, Bowen Yu 0002, Gao Huang 0001, Junyang Lin
NeurIPS15
2025 Anatomical prior-based vertebral landmark detection for spinal disorder diagnosis
Yukang Yang, Ming Sun 0019, Shiji Song, Gao Huang 0001
Artif. Intell. Medicine6
2025 InfoPro: Locally Supervised Deep Learning by Maximizing Information Propagation
Yulin Wang 0002, Zanlin Ni, Yifan Pu, Cai Zhou, Jixuan Ying, Shiji Song, Gao Huang 0001
Int. J. Comput. Vis.6
2025 Uni-AdaFocus: Spatial-Temporal Dynamic Computation for Video Recognition
abstract
This paper presents a comprehensive exploration of the phenomenon of data redundancy in video understanding, with the aim to improve computational efficiency. Our investigation commences with an examination of spatial redundancy, which refers to the observation that the most informative region in each video frame usually corresponds to a small image patch, whose shape, size and location shift smoothly across frames. Motivated by this phenomenon, we formulate the patch localization problem as a dynamic decision task, and introduce a spatially adaptive video recognition approach, termed AdaFocus. In specific, a lightweight encoder is first employed to quickly process the full video sequence, whose features are then utilized by a policy network to identify the most task-relevant regions. Subsequently, the selected patches are inferred by a high-capacity deep network for the final prediction. The complete model can be trained conveniently in an end-to-end manner. During inference, once the informative patch sequence has been generated, the bulk of computation can be executed in parallel, rendering it efficient on modern GPU devices. Furthermore, we demonstrate that AdaFocus can be easily extended by further considering the temporal and sample- wise redundancies, i.e., allocating the majority of computation to the most task-relevant video frames, and minimizing the computation spent on relatively "easier" videos. Our resulting algorithm, Uni-AdaFocus, establishes a comprehensive framework that seamlessly integrates spatial, temporal, and sample- wise dynamic computation, while it preserves the merits of AdaFocus in terms of efficient end-to-end training and hardware friendliness. In addition, Uni-AdaFocus is general and flexible as it is compatible with off-the-shelf backbone models (e.g., TSM and X3D), which can be readily deployed as our feature extractor, yielding a significantly improved computational efficiency. Empirically, extensive experiments based on seven widely-used benchmark datasets (i.e., ActivityNet, FCVID, Mini-Kinetics, Something-Something V1&V2, Jester, and Kinetics-400) and three real-world application scenarios (i.e., fine-grained diving action classification, Alzheimer's and Parkinson's diseases diagnosis with brain magnetic resonance images (MRI), and violence recognition for online videos) substantiate that Uni-AdaFocus is considerably more efficient than the competitive baselines. Code and pre-trained models are available at https://github.com/blackfeather-wang/AdaFocus and https://github.com/LeapLabTHU/AdaFocusV2.
Yulin Wang 0002, Haoji Zhang 0001, Shiji Song, Chao Deng 0002, Junlan Feng, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Cross-modal adapter for vision-language retrieval
Haojun Jiang, Jianke Zhang, Rui Huang 0012, Chunjiang Ge, Zanlin Ni, Shiji Song, Gao Huang 0001
Pattern Recognit.6
2025 Distributionally Robust Chance-Constrained Line Planning for Railway Systems Under Passenger Demand Uncertainty
abstract
In a railway network, making a line plan is a critical optimization problem that determines the daily transportation capacity of the network, aiming at aligning it with passenger demand while minimizing the operational cost. However, the inherent uncertainty in passenger demand often makes the determined line plan infeasible to cover. Meanwhile, additional adjustments to the line plan due to fluctuating demand in the complex railway system can cause unexpected costs. To obtain a robust line plan, we propose a distributionally robust chance-constrained (DRCC) line planning model based on a type-$\infty $Wasserstein ambiguity set, aiming to generate a line plan that remains feasible with a pre-specified probability while accommodating a given distribution deviation tolerance. We present an equivalent tractable reformulation for the proposed DRCC model by explicitly characterizing the worst-case probability distribution. Furthermore, we develop valid inequalities and a warm start strategy tailored to this model to enhance computational efficiency. The proposed model and solution method are validated through numerical experiments conducted on the Wuhan-Guangzhou high-speed railway corridor. Results demonstrate the effectiveness of the solution acceleration techniques and underscore the advantage of the DRCC model over the robust optimization model and conventional chance-constrained model in reducing operational costs. Note to Practitioners—The motivation for this paper stems from the need to provide a daily line plan for the railway to meet the uncertain passenger demand with a pre-specified probability. We assume that an estimated probability distribution of the demand is available and characterized by finite scenarios with weights, but we do not need the estimation to be perfect/accurate. We propose a distributionally robust chance-constrained line planning model to maintain robustness under the deviation/inaccurate probability distribution. To solve the model efficiently, we develop a warm start method to quickly find a partial initial solution and design valid inequalities where adding them as constraints can shorten the solving procedure by reducing the feasible region of the model’s continuous relaxation. The produced line plan requires less operational cost than the robust optimization model to cover the same amount of demand by the benefits of leveraging the distributional information.
Linyu Liu, Wanlu Yang, Shiji Song
IEEE Trans Autom. Sci. Eng.3
2025 Fast Finite-Time Bipartite Formation With Obstacle Avoidance for Time-Delay Multiagent Systems: Application in Mobile Robot Swarm
abstract
The bipartite formation allows agents to achieve two formation structures in opposite directions and finds wide applications in social as well as natural situations. This article investigates the fast finite-time bipartite formation control problem for time-delay nonlinear multiagent systems operating in an obstacle environment. A fuzzy adaptive formation control strategy is proposed utilizing the leader-following method, under which the desired formation is achieved at a fast convergence rate. Since the practical actuator output is usually limited and susceptible to faults, its antisaturation and fault-tolerance capacities are thus considered in the controller construction. Moreover, an obstacle avoidance mechanism is embedded in the control strategy utilizing the artificial potential field method, ensuring the security operation of the formation. An improved Lyapunov–Krasovskii functional is constructed for the stability analysis, which makes full use of the time-delay information of the system. The fast finite-time convergence of the formation error system and the feasibility of obstacle avoidance behavior are verified with the Lyapunov stability theory and the designed energy function. Finally, the proposed control strategy is applied to the bipartite formation control task of a mobile robot swarm, and sufficient simulation results are presented to demonstrate the effectiveness of the theoretical analysis.
Zhenyu Chang, Guangdeng Zong, Shiji Song, Xudong Zhao 0001
IEEE Trans. Ind. Informatics3
2025 Unified Scheduling Model for High-Speed Train Timetable Optimization and Rescheduling Based on Deep Reinforcement Learning
abstract
Train schedule consists of two major phases, train timetable optimization (TTO) and train timetable rescheduling (TTR), which are interconnected with each other and aim to maintain the safety and punctuality of high-speed railway operations under ideal conditions and unexpected disturbances. However, current preparation and adjustment of train timetables face challenges in real-time responsiveness and poor performance on large-scale instances. To alleviate these problems, we propose a unified scheduling model based on deep reinforcement learning (DRL) for both TTO and TTR problems with similar formulations. The key components of our approach include a state representation utilizing the Markov decision process that captures global train and station characteristics, and a policy network that extracts information from this representation to sequentially construct the train departure order. The main benefits of our framework include adaptability to different stopping plans and delay scenarios, decoupling from the problem size, and ensuring the feasibility of generated schemes. Furthermore, to improve the solution quality, we integrate the learned decision policies with a local search method, enabling the scalability of the model with little additional computation cost. Experiments on extensive TTO and TTR instances of the Beijing-Shanghai high-speed railway line demonstrate the effectiveness and practicality of our approach. Our DRL-based method outperforms all the heuristic rules and commercial solvers without retraining the model on various problem sizes, especially on large-scale cases under limited calculation time.
Wanlu Yang, Linyu Liu, Haofeng Yuan, Shiji Song
IEEE Trans. Intell. Transp. Syst.4
2025 Domain Adaptation via Prompt Learning
abstract
Unsupervised domain adaptation (UDA) aims to adapt models learned from a well-annotated source domain to a target domain, where only unlabeled samples are given. Current UDA approaches learn domain-invariant features by aligning source and target feature spaces through statistical discrepancy minimization or adversarial training. However, these constraints could lead to the distortion of semantic feature structures and loss of class discriminability. In this article, we introduce a novel prompt learning paradigm for UDA, named domain adaptation via prompt learning (DAPrompt). In contrast to prior works, our approach learns the underlying label distribution for target domain rather than aligning domains. The main idea is to embed domain information into prompts, a form of representation generated from natural language, which is then used to perform classification. This domain information is shared only by images from the same domain, thereby dynamically adapting the classifier according to each domain. By adopting this paradigm, we show that our model not only outperforms previous methods on several cross-domain benchmarks but also is very efficient to train and easy to implement.
Chunjiang Ge, Rui Huang 0012, Mixue Xie, Zihang Lai, Shiji Song, Shuang Li 0008, Gao Huang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Local Synchronization for Delayed Complex Dynamical Networks via Self-Triggered Impulsive Control Involving Delays
abstract
This article investigates the local synchronization for delayed complex dynamical networks (CDNs) under self-triggered impulsive control (STIC) approaches involving delays. With the help of Lyapunov-Razumikhin methods and comparison principle, some design criteria of STIC strategies ensuring local synchronization for delayed CDNs with delayed impulses are provided, and Zeno behavior can be avoided. Compared with the existing results on synchronization of CDNs under STIC, in this article, time delays in both continuous and discrete system dynamics are well considered. Moreover, the proposed self-triggered mechanism (STM) is an explicit expression, under which the next triggering instant can be derived directly, with simple structure and easy implementation. Finally, two numerical examples are provided to validate the proposed theoretical criteria.
Xiaodi Li 0001, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.3
2025 Advancing Generalization in PINNs Through Latent-Space Representations
abstract
Physics-informed neural networks (PINNs) have made significant strides in modeling dynamical systems governed by partial differential equations (PDEs). However, their generalization capabilities across varying scenarios remain limited. To overcome this limitation, we propose physics-informed dynamics representation learner (PiDo), a novel physics-informed neural PDE solver designed to generalize effectively across diverse PDE configurations, including varying initial conditions, PDE coefficients, and training-time horizons. PiDo exploits the shared underlying structure of dynamical systems with different properties by projecting PDE solutions into a latent space using auto-decoding. It then learns the dynamics of these latent representations, conditioned on the PDE coefficients. Despite its promise, integrating latent dynamics models within a physics-informed framework poses challenges due to the optimization difficulties associated with physics-informed losses. To address these challenges, we introduce a novel approach that diagnoses and mitigates these issues within the latent space. This strategy employs straightforward yet effective regularization techniques, enhancing both the temporal extrapolation performance and the training stability of PiDo. We validate PiDo on a range of benchmarks, including 1-D combined equations and 2-D Navier-Stokes equations. In addition, we demonstrate the transferability of its learned representations to downstream applications such as long-term integration and inverse problems.
Yifan Pu, Shiji Song, Gao Huang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Decoupled Prioritized Resampling for Offline RL
abstract
Offline reinforcement learning (RL) is challenged by the distributional shift problem. To tackle this issue, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy. However, these constraints are applied equally to well-performing and inferior actions through uniform sampling, which might negatively affect the learned policy. In this article, we propose offline decoupled prioritized resampling (ODPR), which designs specialized priority functions for the suboptimal policy constraint issue in offline RL and employs unique decoupled resampling for training stability. Through theoretical analysis, we show that the distinctive priority functions induce a provable improved behavior policy by modifying the distribution of the original behavior policy, and when constrained to this improved policy, a policy-constrained offline RL algorithm is likely to yield a better solution. We provide two practical implementations to balance computation and performance: one estimates priorities based on a fit value network [advantage-based ODPR (ODPR-A)] and the other utilizes trajectory returns [return-based ODPR (ODPR-R)] for quick computation. As a highly compatible plug-and-play component, ODPR is evaluated with five prevalent offline RL algorithms: behavior cloning (BC), twin delayed deep deterministic policy gradient + BC (TD3 + BC), OnestepRL, conservative Q-learning (CQL), and implicit Q-learning (IQL). Our experiments confirm that both ODPR-A and ODPR-R significantly improve performance across all baseline methods. Moreover, ODPR-A can be effective in some challenging settings, i.e., without trajectory information. Code and pretrained weights are available at https://github.com/yueyang130/ODPR.
Bingyi Kang, Xiao Ma 0006, Qisen Yang, Gao Huang 0001, Shiji Song, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.6
2024 A Reinforcement-Learning-Based Multiple-Column Selection Strategy for Column Generation
abstract
Column generation (CG) is one of the most successful approaches for solving large-scale linear programming (LP) problems. Given an LP with a prohibitively large number of variables (i.e., columns), the idea of CG is to explicitly consider only a subset of columns and iteratively add potential columns to improve the objective value. While adding the column with the most negative reduced cost can guarantee the convergence of CG, it has been shown that adding multiple columns per iteration rather than a single column can lead to faster convergence. However, it remains a challenge to design a multiple-column selection strategy to select the most promising columns from a large number of candidate columns. In this paper, we propose a novel reinforcement-learning-based (RL) multiple-column selection strategy. To the best of our knowledge, it is the first RL-based multiple-column selection strategy for CG. The effectiveness of our approach is evaluated on two sets of problems: the cutting stock problem and the graph coloring problem. Compared to several widely used single-column and multiple-column selection strategies, our RL-based multiple-column selection strategy leads to faster convergence and achieves remarkable reductions in the number of CG iterations and runtime.
Haofeng Yuan, Lichang Fang, Shiji Song
AAAI3
2024 PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM Agents
abstract
Qisen Yang, Zekun Wang, Honghui Chen, Shenzhi Wang, Yifan Pu, Xin Gao, Wenhao Huang, Shiji Song, Gao Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Qisen Yang, Honghui Chen, Shenzhi Wang, Yifan Pu, Wenhao Huang 0001, Shiji Song, Gao Huang 0001
ACL (1)8
2024 Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models
abstract
Recently, diffusion models have made remarkable progress in text-to-image (T2I) generation, synthesizing images with highfidelity and diverse contents. Despite this advancement, latent space smoothness within diffusion models remains largely unexplored. Smooth latent spaces en-sure that a perturbation on an input latent corresponds to a steady change in the output image. This property proves beneficial in downstream tasks, including image interpolation, inversion, and editing. In this work, we expose the non-smoothness of diffusion latent spaces by observing noticeable visual fluctuations resulting from minor latent variations. To tackle this issue, we propose Smooth Diffusion, a new category of diffusion models that can be simultaneously high-performing and smooth. Specifically, we introduce Step-wise Variation Regularization to enforce the proportion between the variations of an arbitrary input latent and that of the output image is a constant at any diffusion training step. In addition, we devise an interpolation standard deviation (ISTD) metric to effectively assess the latent space smoothness of a diffusion model. Extensive quantitative and qualitative experiments demonstrate that Smooth Diffusion stands out as a more desirable solution not only in T2I generation but also across various downstream tasks. Smooth Diffusion is implemented as a plug-and-play Smooth-LoRA to work with various community models. Code is available at https://github.com/SHI-Labs/Smooth-Diffusion.
Xingqian Xu, Yifan Pu, Zanlin Ni, Chaofei Wang, Manushree Vasu, Shiji Song, Gao Huang 0001, Humphrey Shi
CVPR7
2024 Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
abstract
The field of image synthesis is currently flourishing due to the advancements in diffusion models. While diffusion models have been successful, their computational inten-sity has prompted the pursuit of more efficient alternatives. As a representative work, non-autoregressive Transformers (NATs) have been recognized for their rapid generation. However, a major drawback of these models is their in-ferior performance compared to diffusion models. In this paper, we aim to re-evaluate the full potential of NATs by revisiting the design of their training and inference strategies. Specifically, we identify the complexities in properly configuring these strategies and indicate the possible sub-optimality in existing heuristic-driven designs. Recognizing this, we propose to go beyond existing methods by directly solving the optimal strategies in an automatic framework. The resulting method, named AutoNAT, advances the performance boundaries of NATs notably, and is able to perform comparably with the latest diffusion models with a significantly reduced inference cost. The effectiveness of AutoNAT is comprehensively validated on four benchmark datasets, i.e., ImageNet-256 & 512, MS-COCO, and CC3M. Code and pretrained models will be available at htt P s: / /gi thub. com/LeapLabTHU/ImprovedNAT.
Zanlin Ni, Yulin Wang 0002, Renping Zhou, Jinyi Hu, Zhiyuan Liu 0001, Shiji Song, Yuan Yao 0013, Gao Huang 0001
CVPR7
2024 GSVA: Generalized Segmentation via Multimodal Large Language Models
abstract
Generalized Referring Expression Segmentation (GRES) extends the scope of classic RES to refer to multiple ob-jects in one expression or identify the empty targets absent in the image. GRES poses challenges in modeling the com-plex spatial relationships of the instances in the image and identifying non-existing referents. Multimodal Large Language Models (MLLMs) have recently shown tremendous progress in these complicated vision-language tasks. Con-necting Large Language Models (LLMs) and vision models, MLLMs are proficient in understanding contexts with visual inputs. Among them, LISA, as a representative, adopts a special [SEG] token to prompt a segmentation mask de-coder, e.g., SAM, to enable MLLMs in the RES task. How-ever, existing solutions to GRES remain unsatisfactory since current segmentation MLLMs cannot correctly handle the cases where users might reference multiple subjects in a singular prompt or provide descriptions incongruent with any image target. In this paper, we propose Generalized Segmentation Vision Assistant (GSVA) to address this gap. Specifically, GSVA reuses the [SEG] token to prompt the segmentation model towards supporting multiple mask ref-erences simultaneously and innovatively learns to generate a [REJ] token to reject the null targets explicitly. Ex-periments validate GSVA's efficacy in resolving the GRES issue, marking a notable enhancement and setting a new record on the GRES benchmark gRefCOCO dataset. GSVA also proves effective across various classic referring seg-mentation and comprehension tasks. Code is available at https://github.com/LeapLabTHU/GSVA.
Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, Gao Huang 0001
CVPR5
2024 Agent Attention: On the Integration of Softmax and Linear Attention
Dongchen Han, Tianzhu Ye, Yizeng Han, Zhuofan Xia, Siyuan Pan, Pengfei Wan 0001, Shiji Song, Gao Huang 0001
ECCV (50)7
2024 Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation Without Manual Labels
Rui Huang 0012, Songyou Peng, Ayça Takmaz, Federico Tombari, Marc Pollefeys, Shiji Song, Gao Huang 0001, Francis Engelmann
ECCV (34)6
2024 Efficient Diffusion Transformer with Step-Wise Dynamic Attention Mediators
Yifan Pu, Zhuofan Xia, Dongchen Han, Qixiu Li, Yuhui Yuan, Ji Li 0006, Yizeng Han, Shiji Song, Gao Huang 0001, Xiu Li 0001
ECCV (15)10
2024 DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
Le Yang 0007, Ziwei Zheng, Yizeng Han, Hao Cheng 0015, Shiji Song, Gao Huang 0001, Fan Li 0003
ECCV (46)5
2024 Visual Attention Based Cognitive Human-Robot Collaboration for Pedicle Screw Placement in Robot-Assisted Orthopedic Surgery
abstract
Current orthopedic robotic systems largely focus on navigation, aiding surgeons in positioning a guiding tube but still requiring manual drilling and screw placement. The automation of this task not only demands high precision and safety due to the intricate physical interactions between the surgical tool and bone but also poses significant risks when executed without adequate human oversight. As it involves continuous physical interaction, the robot should collaborate with the surgeon, understand the human intent, and always include the surgeon in the loop. To achieve this, this paper proposes a new cognitive human–robot collaboration framework, including the intuitive AR-haptic human–robot interface, the visual-attention-based surgeon model, and the shared interaction control scheme for the robot. User studies on a robotic platform for orthopedic surgery are presented to illustrate the performance of the proposed method. The results demonstrate that the proposed human– robot collaboration framework outperforms full robot and full human control in terms of safety and ergonomics.
Chen Chen 0087, Qikai Zou, Yuhang Song 0009, Mingrui Yu 0001, Senqiang Zhu, Shiji Song, Xiang Li 0009
IROS6
2024 A Unified Interaction Control Framework for Safe Robotic Ultrasound Scanning with Human-Intention-Aware Compliance
abstract
The ultrasound scanning robot operates in environments where frequent human-robot interactions occur. Most existing control methods for ultrasound scanning address only one specific interaction situation or implement hard switches between controllers for different situations, which compromises both safety and efficiency. In this paper, we propose a unified interaction control framework for ultrasound scanning robots capable of handling all common interactions, distinguishing both human-intended and unintended types, and adapting with appropriate compliance. Specifically, the robot suspends or modulates its ongoing main task if the interaction is intended, e.g., when the doctor grasps the robot to lead the end effector actively. Furthermore, it can identify unintended interactions and avoid potential collision in the null space beforehand. Even if that collision has happened, it can become compliant with the collision in the null space and try to reduce its impact on the main task (where the scan is ongoing) kinematically and dynamically. The multiple situations are integrated into a unified controller with a smooth transition to deal with the interactions by exhibiting human-intention-aware compliance. Experimental results validate the framework’s ability to cope with all common interactions including intended intervention and unintended collision in a collaborative carotid artery ultrasound scanning task.
Xiangjie Yan, Shaqi Luo, Yongpeng Jiang, Mingrui Yu 0001, Chen Chen 0087, Senqiang Zhu, Gao Huang 0001, Shiji Song, Xiang Li 0009
IROS8
2024 In-Hand Following of Deformable Linear Objects Using Dexterous Fingers with Tactile Sensing
abstract
Most research on deformable linear object (DLO) manipulation assumes rigid grasping. However, beyond rigid grasping and re-grasping, in-hand following is also an essential skill that humans use to dexterously manipulate DLOs, which requires continuously changing the grasp point by in-hand sliding while holding the DLO to prevent it from falling. Achieving such a skill is very challenging for robots without using specially designed but not versatile end-effectors. Previous works have attempted using generic parallel grippers, but their robustness is unsatisfactory owing to the conflict between following and holding, which is hard to balance with a one-degree-of-freedom gripper. In this work, inspired by how humans use fingers to follow DLOs, we explore the usage of a generic dexterous hand with tactile sensing to imitate human skills and achieve robust in-hand DLO following. To enable the hardware system to function in the real world, we develop a framework that includes Cartesian-space arm-hand control, tactile-based in-hand 3-D DLO pose estimation, and task-specific motion design. Experimental results demonstrate the significant superiority of our method over using parallel grippers, as well as its great robustness, generalizability, and efficiency.
Mingrui Yu 0001, Boyuan Liang, Xiang Zhang 0020, Xinghao Zhu, Lingfeng Sun, Shiji Song, Xiang Li 0009, Masayoshi Tomizuka
IROS7
2024 Cardiac Copilot: Automatic Probe Guidance for Echocardiography with World Model
Haojun Jiang, Zhenguo Sun, Meng Li 0087, Yu Sun 0020, Shaqi Luo, Shiji Song, Gao Huang 0001
MICCAI (1)7
2024 Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
Ziwei Zheng, Zechuan Zhang, Yulin Wang 0002, Shiji Song, Gao Huang 0001, Le Yang 0007
ACM Multimedia4
2024 Bridging the Divide: Reconsidering Softmax and Linear Attention
abstract
Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when dealing with high-resolution inputs. In contrast, linear attention naturally enjoys linear complexity and has great potential to scale up to higher-resolution images. Nonetheless, the unsatisfactory performance of linear attention greatly limits its practical application in various scenarios. In this paper, we take a step forward to close the gap between the linear and Softmax attention with novel theoretical analyses, which demystify the core factors behind the performance deviations. Specifically, we present two key perspectives to understand and alleviate the limitations of linear attention: the injective property and the local modeling ability. Firstly, we prove that linear attention is not injective, which is prone to assign identical attention weights to different query vectors, thus adding to severe semantic confusion since different queries correspond to the same outputs. Secondly, we confirm that effective local modeling is essential for the success of Softmax attention, in which linear attention falls short. The aforementioned two fundamental differences significantly contribute to the disparities between these two attention paradigms, which is demonstrated by our substantial empirical validation in the paper. In addition, more experiment results indicate that linear attention, as long as endowed with these two properties, can outperform Softmax attention across various tasks while maintaining lower computation complexity. Code is available at https://github.com/LeapLabTHU/InLine.
Dongchen Han, Yifan Pu, Zhuofan Xia, Yizeng Han, Xuran Pan, Xiu Li 0001, Jiwen Lu, Shiji Song, Gao Huang 0001
NeurIPS8
2024 Demystify Mamba in Vision: A Linear Attention Perspective
abstract
Mamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba model shares surprising similarities with linear attention Transformer, which typically underperform conventional Transformer in practice. By exploring the similarities and disparities between the effective Mamba and subpar linear attention Transformer, we provide comprehensive analyses to demystify the key factors behind Mamba’s success. Specifically, we reformulate the selective state space model and linear attention within a unified formulation, rephrasing Mamba as a variant of linear attention Transformer with six major distinctions: input gate, forget gate, shortcut, no attention normalization, single-head, and modified block design. For each design, we meticulously analyze its pros and cons, and empirically evaluate its impact on model performance in vision tasks. Interestingly, the results highlight the forget gate and block design as the core contributors to Mamba’s success, while the other four designs are less crucial. Based on these findings, we propose a Mamba- Inspired Linear Attention (MILA) model by incorporating the merits of these two key designs into linear attention. The resulting model outperforms various vision Mamba models in both image classification and high-resolution dense prediction tasks, while enjoying parallelizable computation and fast inference speed. Code is available at https://github.com/LeapLabTHU/MLLA.
Dongchen Han, Zhuofan Xia, Yizeng Han, Yifan Pu, Chunjiang Ge, Shiji Song, Bo Zheng 0007, Gao Huang 0001
NeurIPS8
2024 DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable comprehension and reasoning capabilities with complex language and visual data. These advances have spurred the vision of establishing a generalist robotic MLLM proficient in understanding complex human instructions and accomplishing various embodied tasks, whose feasibility has been recently verified~\cite{rt-2,rt-x}. However, developing MLLMs for real-world robots is challenging due to the typically limited computation and memory capacities available on robotic platforms. In contrast, the inference of MLLMs usually incorporates storing billions of parameters and performing tremendous computation, imposing significant hardware demands. In our paper, we seek to address this challenge by leveraging an intriguing observation: relatively easier situations make up the bulk of the procedure of controlling robots to fulfill diverse tasks, and they generally require far smaller models to obtain the correct robotic actions. Motivated by this observation, we propose a \emph{Dynamic Early-Exit for Robotic MLLM} (DeeR) framework that automatically adjusts the size of the activated MLLM based on each situation at hand. The approach leverages a multi-exit architecture in MLLMs, which allows the model to cease processing once a proper size of the model has been activated for a specific situation, thus avoiding further redundant computation. Additionally, we develop novel algorithms that establish early-termination criteria for DeeR, conditioned on predefined demands such as average computational cost (\emph{i.e.}, power consumption), as well as peak computational consumption (\emph{i.e.}, latency) and GPU memory usage. These enhancements ensure that DeeR operates efficiently under varying resource constraints while maintaining competitive performance. Moreover, we design a tailored training method for integrating temporal information on top of such multi-exit architectures to predict actions reasonably. On the CALVIN robot manipulation benchmark, DeeR demonstrates significant reductions in computational costs by 5.2-6.5x and GPU memory by 2x without compromising performance. Code and checkpoints are available at https://github.com/yueyang130/DeeR-VLA.
Yulin Wang 0002, Bingyi Kang, Yizeng Han, Shenzhi Wang, Shiji Song, Jiashi Feng, Gao Huang 0001
NeurIPS6
2024 A new artificial bee colony algorithm for the flexible job shop scheduling problem with extra resource constraints in numeric control centers
Xiaoya Liao, Rui Zhang 0039, Yali Chen 0004, Shiji Song
Expert Syst. Appl.4
2024 Solving flexible job shop scheduling problems via deep reinforcement learning
Erdong Yuan, Shuli Cheng, Shiji Song
Expert Syst. Appl.4
2024 Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
abstract
Long-tailed distributions frequently emerge in real-world data, where a large number of minority categories contain a limited number of samples. Such imbalance issue considerably impairs the performance of standard supervised learning algorithms, which are mainly designed for balanced training sets. Recent investigations have revealed that supervised contrastive learning exhibits promising potential in alleviating the data imbalance. However, the performance of supervised contrastive learning is plagued by an inherent challenge: it necessitates sufficiently large batches of training data to construct contrastive pairs that cover all categories, yet this requirement is difficult to meet in the context of class-imbalanced data. To overcome this obstacle, we propose a novel probabilistic contrastive (ProCo) learning algorithm that estimates the data distribution of the samples from each class in the feature space, and samples contrastive pairs accordingly. In fact, estimating the distributions of all classes using features in a small batch, particularly for imbalanced data, is not feasible. Our key idea is to introduce a reasonable and simple assumption that the normalized features in contrastive learning follow a mixture of von Mises-Fisher (vMF) distributions on unit space, which brings two-fold benefits. First, the distribution parameters can be estimated using only the first sample moment, which can be efficiently computed in an online manner across different batches. Second, based on the estimated distribution, the vMF distribution allows us to sample an infinite number of contrastive pairs and derive a closed form of the expected contrastive loss for efficient optimization. Other than long-tailed problems, ProCo can be directly applied to semi-supervised learning by generating pseudo-labels for unlabeled data, which can subsequently be utilized to estimate the distribution of the samples inversely. Theoretically, we analyze the error bound of ProCo. Empirically, extensive experimental results on supervised/semi-supervised visual recognition and object detection tasks demonstrate that ProCo consistently outperforms existing methods across various datasets.
Chaoqun Du, Yulin Wang 0002, Shiji Song, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Latency-Aware Unified Dynamic Networks for Efficient Image Recognition
abstract
Dynamic networks have become a pivotal area of study in deep learning due to their ability to selectively activate computing units (such as layers or channels) or dynamically allocate computation to information-rich regions. This capability significantly curtails unnecessary computations, adapting to varying inputs. Despite these advantages, the practical efficiency of dynamic models often falls short of theoretical computation. This discrepancy arises from three primary challenges: 1) a lack of a unified framework across different dynamic inference paradigms due to the fragmented research landscape; 2) an excessive focus on algorithm design at the expense of scheduling strategies, which are essential for optimizing resource utilization on hardware; and 3) the complexity of latency evaluation, since most current libraries cater to static operators. To tackle these issues, we introduce Latency-Aware Unified Dynamic Networks (LAUDNet), a general framework that integrates three fundamental dynamic paradigms-spatially-adaptive computation, layer skipping, and channel skipping-into a single unified formulation. LAUDNet not only refines algorithmic design but also enhances scheduling optimization with the aid of a latency predictor. This predictor efficiently and accurately predicts the inference latency of dynamic operators on specific hardware setups. Our empirical assessments across multiple vision tasks-image classification, object detection, and instance segmentation-confirm that LAUDNet significantly bridges the gap between theoretical and practical efficiency. For instance, LAUDNet cuts down the practical latency of its static counterpart, ResNet-101, by over 50% on hardware platforms like V100, RTX 3090, and TX2 GPUs. Additionally, LAUDNet excels in the accuracy-efficiency trade-off compared to other methods.
Yizeng Han, Zhihang Yuan, Yifan Pu, Chaofei Wang, Shiji Song, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
abstract
The superior performance of modern computer vision backbones (e.g., vision Transformers learned on ImageNet-1 K/22 K) usually comes with a costly training procedure. This study contributes to this issue by generalizing the idea of curriculum learning beyond its original formulation, i.e., training models using easier-to-harder data. Specifically, we reformulate the training curriculum as a soft-selection function, which uncovers progressively more difficult patterns within each example during training, instead of performing easier-to-harder sample selection. Our work is inspired by an intriguing observation on the learning dynamics of visual backbones: during the earlier stages of training, the model predominantly learns to recognize some 'easier-to-learn' discriminative patterns in the data. These patterns, when observed through frequency and spatial domains, incorporate lower-frequency components, and the natural image contents without distortion or data augmentation. Motivated by these findings, we propose a curriculum where the model always leverages all the training data at every learning stage, yet the exposure to the 'easier-to-learn' patterns of each example is initiated first, with harder patterns gradually introduced as training progresses. To implement this idea in a computationally efficient way, we introduce a cropping operation in the Fourier spectrum of the inputs, enabling the model to learn from only the lower-frequency components. Then we show that exposing the contents of natural images can be readily achieved by modulating the intensity of data augmentation. Finally, we integrate these two aspects and design curriculum learning schedules by proposing tailored searching algorithms. Moreover, we present useful techniques for deploying our approach efficiently in challenging practical scenarios, such as large-scale parallel training, and limited input/output or data pre-processing speed. The resulting method, EfficientTrain++, is simple, general, yet surprisingly effective. As an off-the-shelf approach, it reduces the training time of various popular models (e.g., ResNet, ConvNeXt, DeiT, PVT, Swin, CSWin, and CAFormer) by [Formula: see text] on ImageNet-1 K/22 K without sacrificing accuracy. It also demonstrates efficacy in self-supervised learning (e.g., MAE).
Yulin Wang 0002, Rui Lu 0001, Yizeng Han, Shiji Song, Gao Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Joint representation learning for text and 3D point cloud
Rui Huang 0012, Xuran Pan, Henry Zheng, Haojun Jiang, Cheng Wu 0002, Shiji Song, Gao Huang 0001
Pattern Recognit.7
2024 A unified framework for convolution-based graph neural networks
Xuran Pan, Xiaoyan Han, Chaofei Wang, Shiji Song, Gao Huang 0001, Cheng Wu 0002
Pattern Recognit.5
2024 Data-Driven Distributionally Robust Optimization for Railway Timetabling Problem
abstract
Unpredictable disturbances that occur when the train is running often make the actual timetable deviate from the planned timetable and deteriorate service quality for passengers. To provide a robust planned timetable that is reliable in operation, this paper presents a two-stage distributionally robust railway timetabling (DRRT) model, where the decisions in the first and second stages are the planned and actual timetables, respectively. To deal with the uncertain travel time, the proposed DRRT model assumes the true distribution belongs to a Wasserstein ambiguity set centered at the empirical distribution constructed from finite samples and aims to minimize the sum of the total idle time, travel time, and worst-case expected delay over the ambiguity set. To solve this model, we design a decomposition algorithm with provable convergence, which iteratively solves a master problem and a set of subproblems. To enhance the algorithm efficiency, a blend of two popular decomposition methods is proposed to integrate their advantages, and the solution of the subproblems is accelerated by using an uncapacitated minimum-cost flow interpretation of the second-stage optimization problem. Experiments on real-world instances demonstrate the effectiveness of these enhancements. The results also show that our model prominently outperforms the sample average approximation model in terms of the out-of-sample performance, especially when only a small sample set is available. Moreover, at the expense of equal traffic efficiency, our model is more robust than those models that do not utilize distribution information. Note to Practitioners—This paper was motivated by the problem of using finite historical data to provide a robust planned railway timetable, aiming at minimizing the traffic efficiency losses and expected delays under the worst-case distribution. Existing studies on dealing with the uncertain travel time usually assumed that the distribution is exactly known, or the distribution is completely unknown but the travel time belongs to an uncertainty set. However, in practice, the exact distribution is hard to know but can only be coarsely estimated from data. This paper proposes a data-driven two-stage model by assuming that the true distribution belongs to a 1-Wasserstein ball centered at the empirical distribution with a bounded support set, which is a middle ground between the above two distribution assumptions. To solve this problem efficiently, we propose a customized algorithm with provable convergence and then enhance it by exploiting problem-specific properties. This distributionally robust model can be easily shifted to other timetabling problems, and the algorithm enhancement techniques can also be applied to efficiently solve other two-stage recourse models with similar properties.
Linyu Liu, Shiji Song, Zhuolin Wang
IEEE Trans Autom. Sci. Eng.2
2024 A Review of Robust Machine Scheduling
abstract
Robust optimization (RO) has been recognized as an effective means to deal with unanticipated events in highly uncertain and risky environments. This paper systematically reviews two types of emerging RO machine scheduling approaches—robust machine scheduling (R-MS) and distributionally R-MS (DR-MS) methods—which usually offer tractable formulations and analytical results for machine scheduling problems under uncertainty. First, after highlighting the advantages of RO methods over the stochastic approach in terms of tractability and robustness, we use the bibliometric method to analyze the literature related to R-MS/DR-MS problems and classify them from the following aspects: (1) uncertain factors, (2) uncertainty descriptions, (3) robustness criteria, (4) machine environments and (5) solution methods. Second, we discuss the uncertainty descriptions, and the robust feasibility and robust optimality criteria. We further provide a state-of-the-art review of R-MS/DR-MS models in different machine environments and discuss the performance of the R-MS/DR-MS models. Third, we review and discuss the existing exact, approximation, online, and heuristic solution methods for solving R-MS/DR-MS models. Finally, we present future research opportunities in two promising areas: green machine scheduling problems and machine learning-enabled algorithms.Note to Practitioners—Machine scheduling plays an essential role in industrial and service systems, such as manufacturing, power generation, transportation and medical systems. However, in practice, scheduling systems usually operate in highly uncertain environments due to noisy measurements, prediction errors, and implementation deviations. To ensure robust feasibility and robust optimality, robust machine scheduling (R-MS) and distributionally R-MS (DR-MS) approaches have been recently proposed to hedge against the uncertainties related to processing time, release time, due date, machine breakdown, etc. This paper provides a comprehensive review of the R-MS/DR-MS models and algorithms in different machine environments from the aspects of uncertainty descriptions, robustness criteria and solution methods. This paper further highlights the challenges of R-MS problems and provides promising and valuable research opportunities in terms of problem formulations and algorithm designs.
Ningwei Zhang, Shiji Song, C. L. Philip Chen
IEEE Trans Autom. Sci. Eng.3
2024 FaceCLIP: Facial Image-to-Video Translation via a Brief Text Description
abstract
The existing image-to-video translation methods generally follow a frame-by-frame generative paradigm, while extracting the temporal information from a reference video or an audio stream. Inspired by the recent success in text-guided image generation, we explore a more challenging but promising task, Text-guided Image-to-Video (TI2V) translation. Given an image and a brief text description as input, TI2V aims to generate a facial expression video following the image and text. To this end, we first propose an automatic video captioning pipeline to generate dense textual descriptions for facial video datasets, using both expression labels and action units. These dense textual descriptions provide precise semantic guidance for TI2V learning. Then we design and train an efficient framework, FaceCLIP, on these datasets to deal with the TI2V translation task. FaceCLIP adopts a video autoencoder to model the temporal information of training videos, and a pretrained CLIP model to embed the video frames and the text description. We design a reconstruction loss and an embedding alignment loss to train the autoencoder to obtain the text-guided video generative ability. Recognizing that expressions are closely tied to facial landmark motions, the reconstruction loss is applied to facial landmarks rather than each video frame, significantly enhancing training efficiency. We compare FaceCLIP with several potential baseline methods, and extensively evaluate the performance using multiple metrics. Both qualitative and quantitative results validate the superiority of FaceCLIP in terms of both visual quality and expression-text consistency. Moreover, the unique ability of FaceCLIP to generate videos based on abstract texts demonstrates its stronger generalization capability.
Hayk Manukyan 0001, Chaofei Wang, Levon Khachatryan, Shant Navasardyan, Shiji Song, Humphrey Shi, Gao Huang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 OStr-DARTS: Differentiable Neural Architecture Search Based on Operation Strength
abstract
Differentiable architecture search (DARTS) has emerged as a promising technique for effective neural architecture search, and it mainly contains two steps to find the high-performance architecture. First, the DARTS supernet that consists of mixed operations will be optimized via gradient descent. Second, the final architecture will be built by the selected operations that contribute the most to the supernet. Although DARTS improves the efficiency of neural architecture search (NAS), it suffers from the well-known degeneration issue which can lead to deteriorating architectures. Existing works mainly attribute the degeneration issue to the failure of its supernet optimization, while little attention has been paid to the selection method. In this article, we cease to apply the widely-used magnitude-based selection method and propose a novel criterion based on operation strength that estimates the importance of an operation by its effect on the final loss. We show that the degeneration issue can be effectively addressed by using the proposed criterion without any modification of supernet optimization, indicating that the magnitude-based selection method can be a critical reason for the instability of DARTS. The experiments on NAS-Bench-201 and DARTS search spaces show the effectiveness of our method.
Le Yang 0007, Ziwei Zheng, Yizeng Han, Shiji Song, Gao Huang 0001, Fan Li 0003
IEEE Trans. Cybern.4
2024 Unraveling the Accuracy Enigma: Investigating ZTD Data Precision in TUW-VMF3 and GFZ-VMF3 Products Using a Comprehensive Global GPS Dataset
abstract
Tropospheric delay is one of the major error sources for space geodetic techniques, such as the Global Navigation Satellite Systems (GNSS). The accurate priori zenith tropospheric delay (ZTD) information is of crucial importance for the accuracy of positioning and navigation and reducing the convergence time in GNSS data processing. This paper conducted the first globally spatiotemporal assessment of two existing tropospheric products from Vienna University of Technology (TU Wien) and the GeoForschungsZentrum Potsdam (GFZ), namely TUW-VMF3 and GFZ-VMF3. The referenced ZTD data used in this study are sourced from a screened ZTD dataset created by the Karlsruhe Institute of Technology (KIT) team in 2020, which obtained 91,088,258 screened ZTD values from 12,552 GNSS stations. The results revealed that the two products exhibited high similarity with RMSE/Bias of 16.47/-3.62 and 17.63/-2.23 mm for the TUW-VMF3 and GFZ-VMF3, respectively. Their different performances were also clearly observed from the analysis of the geographical distribution, season and epochs. The TUW-VMF3 outperformed GFZ-VMF3 at 70.1% stations, but noted that in the European region, the GFZ-VMF3 was superior to the TUM-VMF3. In the regions of the North America and Japan, both products exhibited better performance in winter, but in other regions, this seasonal trend was less pronounced. TUW-VMF3 performed better in case #1 without interpolation for two products, and GFZ-VMF3 performed better in case #2, where only TUM-VMF3 required interpolation. In case #3, where both required interpolation, TUW-VMF3 exhibited better with an RMSE of 15.59 mm, although GFZ-VMF3 outperformed in seven epochs.
Debao Yuan, Shiji Song, Juntao Tan, Zhuoyue Wen
IEEE Trans. Geosci. Remote. Sens.5
2024 Hundreds Guide Millions: Adaptive Offline Reinforcement Learning With Expert Guidance
abstract
Offline reinforcement learning (RL) optimizes the policy on a previously collected dataset without any interactions with the environment, yet usually suffers from the distributional shift problem. To mitigate this issue, a typical solution is to impose a policy constraint on a policy improvement objective. However, existing methods generally adopt a "one-size-fits-all" practice, i.e., keeping only a single improvement-constraint balance for all the samples in a mini-batch or even the entire offline dataset. In this work, we argue that different samples should be treated with different policy constraint intensities. Based on this idea, a novel plug-in approach named guided offline RL (GORL) is proposed. GORL employs a guiding network, along with only a few expert demonstrations, to adaptively determine the relative importance of the policy improvement and policy constraint for every sample. We theoretically prove that the guidance provided by our method is rational and near-optimal. Extensive experiments on various environments suggest that GORL can be easily installed on most offline RL algorithms with statistically significant performance improvements.
Qisen Yang, Shenzhi Wang, Qihang Zhang, Gao Huang 0001, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.5
2024 Learning to Assist Different Wearers in Multitasks: Efficient and Individualized Human-in-the-Loop Adaptation Framework for Lower-Limb Exoskeleton
abstract
One of the typical purposes of using lower-limb exoskeleton robots is to provide assistance to the wearer by supporting their weight and augmenting their physical capabilities according to a given task and human motion intentions. The generalizability of robots across different wearers in multiple tasks is important to ensure that the robot can provide correct and effective assistance in actual implementation. However, most lower-limb exoskeleton robots exhibit only limited generalizability. Therefore, this article proposes a human-in-the-loop learning and adaptation framework for exoskeleton robots to improve their performance in various tasks and for different wearers. To suit different wearers, an individualized walking trajectory is generated online using dynamic movement primitives and Bayes optimization. To accommodate various tasks, a task translator is constructed using a neural network to generalize a trajectory to more complex scenarios. These generalization techniques are integrated into a unified variable impedance model, which regulates the exoskeleton to provide assistance while ensuring safety. In addition, an anomaly detection network is developed to quantitatively evaluate the wearer's comfort, which is considered in the trajectory learning procedure and contributes to the relaxation of conflicts in impedance control. The proposed framework is easy to implement, because it requires proprioceptive sensors only to perform and deploy data-efficient learning schemes. This makes the exoskeleton practical for deployment in complex scenarios, accommodating different walking patterns, habits, tasks, and conflicts. Experiments and comparative studies on a lower-limb exoskeleton robot are performed to demonstrate the effectiveness of the proposed framework.
Shu Miao, Gong Chen 0001, Jing Ye 0005, Chenglong Fu 0001, Bin Liang 0001, Shiji Song, Xiang Li 0009
IEEE Trans. Robotics7
2024 A Self-Triggered Impulsive Approach to Group Consensus of MASs With Sensing/Actuation Delays
abstract
This article presents a self-triggered impulsive framework for group consensus of multiagent systems (MASs). Two types of self-triggered delayed impulsive control schemes are proposed to regulate impulsive protocols with sensing and actuation delays, respectively. Here, the Lyapunov-based and comparison-system-based approaches are constructed to achieve the iterative updates of impulse sequences with flexibility, especially the upper bound or average interval of impulsive periods is not restricted explicitly. In addition, several sufficient criteria for multigroup consensus of MASs with sensing and actuation delays are presented, where the correlation inequalities between trigger parameters, time delays, and control strengths are established to promote the co-design of impulsive controller and self-triggering algorithm. The Zeno behavior could be successfully eliminated. It is shown that the presented self-triggered schemes do not necessitate continuous or periodic event-detections and the interaction for neighboring agents works in an impulsive manner, which significantly saves the resource consumption of communication and control. Finally, two numerical examples illustrate the effectiveness of the proposed schemes.
Xiaodi Li 0001, Shiji Song, Xinzhi Liu
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Leveraging Reward Consistency for Interpretable Feature Discovery in Reinforcement Learning
abstract
The black-box nature of deep reinforcement learning (RL) hinders them from real-world applications. Therefore, interpreting and explaining RL agents have been active research topics in recent years. Existing methods for post-hoc explanations usually adopt the action matching principle to enable an easy understanding of vision-based RL agents. In this article, it is argued that the commonly used action matching principle is more like an explanation of deep neural networks (DNNs) than the interpretation of RL agents. It may lead to irrelevant or misplaced feature attribution when different DNNs’ outputs lead to the same rewards or different rewards result from the same outputs. Therefore, we propose to consider rewards, the essential objective of RL agents, as the essential objective of interpreting RL agents as well. To ensure reward consistency during interpretable feature discovery, a novel framework (RL interpreting RL, denoted as RL-in-RL) is proposed to solve the gradient disconnection from actions to rewards. We verify and evaluate our method on the Atari 2600 games as well as Duckietown, a challenging self-driving car simulator environment. The results show that our method manages to keep reward (or return) consistency and achieves high-quality feature attribution. Further, a series of analytical experiments validate our assumption of the action matching principle’s limitations.
Qisen Yang, Huanqian Wang, Mukun Tong, Wenjie Shi, Gao Huang 0001, Shiji Song
IEEE Trans. Syst. Man Cybern. Syst.6
2023 Causal Intervention for Human Trajectory Prediction with Cross Attention Mechanism
abstract
Human trajectory Prediction (HTP) in complex social environments plays a crucial and fundamental role in artificial intelligence systems. Conventional methods make use of both history behaviors and social interactions to forecast future trajectories. However, we demonstrate that the social environment is a confounder that misleads the model to learn spurious correlations between history and future trajectories. To end this, we first formulate the social environment, history and future trajectory variables into a structural causal model to analyze the causalities among them. Based on causal intervention rather than conventional likelihood, we propose a Social Environment ADjustment (SEAD) method, to remove the confounding effect of the social environment. The core of our method is implemented by a Social Cross Attention (SCA) module, which is universal, simple and effective. Our method has consistent improvements on ETH-UCY datasets with three baseline models and achieves competitive performances with existing methods.
Chunjiang Ge, Shiji Song, Gao Huang 0001
AAAI2
2023 Zero-Shot Generative Model Adaptation via Image-Specific Prompt Learning
abstract
Recently, CLIP-guided image synthesis has shown appealing performance on adapting a pre-trained source-domain generator to an unseen target domain. It does not require any target-domain samples but only the textual domain labels. The training is highly efficient, e.g., a few minutes. However, existing methods still have some limitations in the quality of generated images and may suffer from the mode collapse issue. A key reason is that a fixed adaptation direction is applied for all cross-domain image pairs, which leads to identical supervision signals. To address this issue, we propose an Image-specific Prompt Learning (IPL) method, which learns specific prompt vectors for each source-domain image. This produces a more precise adaptation direction for every cross-domain image pair, endowing the target-domain generator with greatly enhanced flexibility. Qualitative and quantitative evaluations on various domains demonstrate that IPL effectively improves the quality and diversity of synthesized images and alleviates the mode collapse. Moreover, IPL is independent of the structure of the generative model, such as generative adversarial networks or diffusion models. Code is available at https://github.com/Picsart-AI-Research/IPL-Zero-Shot-Generative-Model-Adaptation.
Chaofei Wang, Eric J. Zhang, Kai Wang 0058, Xingqian Xu, Shiji Song, Humphrey Shi, Gao Huang 0001
CVPR7
2023 Slide-Transformer: Hierarchical Vision Transformer with Local Self-Attention
abstract
Self-attention mechanism has been a key factor in the recent progress of Vision Transformer (ViT), which enables adaptive feature extraction from global contexts. However, existing self-attention methods either adopt sparse global attention or window attention to reduce the computation complexity, which may compromise the local feature learning or subject to some handcrafted designs. In contrast, local attention, which restricts the receptive field of each query to its own neighboring pixels, enjoys the benefits of both convolution and self-attention, namely local inductive bias and dynamic feature selection. Nevertheless, current local attention modules either use inefficient Im2Col function or rely on specific CUDA kernels that are hard to generalize to devices without CUDA support. In this paper, we propose a novel local attention module, Slide Attention, which leverages common convolution operations to achieve high efficiency, flexibility and generalizability. Specifically, we first re-interpret the column-based Im2Col function from a new row-based perspective and use Depthwise Convolution as an efficient substitution. On this basis, we propose a deformed shifting module based on the re-parameterization technique, which further relaxes the fixed key/value positions to deformed features in the local region. In this way, our module realizes the local attention paradigm in both efficient and flexible manner. Extensive experiments show that our slide attention module is applicable to a variety of advanced Vision Transformer models and compatible with various hardware devices, and achieves consistently improved performances on comprehensive benchmarks.
Xuran Pan, Tianzhu Ye, Zhuofan Xia, Shiji Song, Gao Huang 0001
CVPR4
2023 Dynamic Perceiver for Efficient Visual Recognition
abstract
Early exiting has become a promising approach to improving the inference efficiency of deep networks. By structuring models with multiple classifiers (exits), predictions for "easy" samples can be generated at earlier exits, negating the need for executing deeper layers. Current multi-exit networks typically implement linear classifiers at intermediate layers, compelling low-level features to encapsulate high-level semantics. This sub-optimal design invariably undermines the performance of later exits. In this paper, we propose Dynamic Perceiver (Dyn-Perceiver) to decouple the feature extraction procedure and the early classification task with a novel dual-branch architecture. A feature branch serves to extract image features, while a classification branch processes a latent code assigned for classification tasks. Bi-directional cross-attention layers are established to progressively fuse the information of both branches. Early exits are placed exclusively within the classification branch, thus eliminating the need for linear separability in low-level features. Dyn-Perceiver constitutes a versatile and adaptable framework that can be built upon various architectures. Experiments on image classification, action recognition, and object detection demonstrate that our method significantly improves the inference efficiency of different backbones, outperforming numerous competitive approaches across a broad range of computational budgets. Evaluation on both CPU and GPU platforms substantiate the superior practical efficiency of Dyn-Perceiver. Code is available at https://www.github.com/LeapLabTHU/Dynamic_Perceiver.
Yizeng Han, Dongchen Han, Yulin Wang 0002, Xuran Pan, Yifan Pu, Chao Deng 0002, Junlan Feng, Shiji Song, Gao Huang 0001
ICCV9
2023 FLatten Transformer: Vision Transformer using Focused Linear Attention
abstract
The quadratic computation complexity of self-attention has been a persistent challenge when applying Transformer models to vision tasks. Linear attention, on the other hand, offers a much more efficient alternative with its linear complexity by approximating the Softmax operation through carefully designed mapping functions. However, current linear attention approaches either suffer from significant performance degradation or introduce additional computation overhead from the mapping functions. In this paper, we propose a novel Focused Linear Attention module to achieve both high efficiency and expressiveness. Specifically, we first analyze the factors contributing to the performance degradation of linear attention from two perspectives: the focus ability and feature diversity. To overcome these limitations, we introduce a simple yet effective mapping function and an efficient rank restoration module to enhance the expressiveness of self-attention while maintaining low computation complexity. Extensive experiments show that our linear attention module is applicable to a variety of advanced vision Transformers, and achieves consistently improved performances on multiple benchmarks. Code is available at https://github.com/LeapLabTHU/FLatten-Transformer.
Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, Gao Huang 0001
ICCV4
2023 Adaptive Rotated Convolution for Rotated Object Detection
abstract
Rotated object detection aims to identify and locate objects in images with arbitrary orientation. In this scenario, the oriented directions of objects vary considerably across different images, while multiple orientations of objects exist within an image. This intrinsic characteristic makes it challenging for standard backbone networks to extract high-quality features of these arbitrarily orientated objects. In this paper, we present Adaptive Rotated Convolution (ARC) module to handle the afore-mentioned challenges. In our ARC module, the convolution kernels rotate adaptively to extract object features with varying orientations in different images, and an efficient conditional computation mechanism is introduced to accommodate the large orientation variations of objects within an image. The two designs work seamlessly in rotated object detection problem. Moreover, ARC can conveniently serve as a plug-and-play module in various vision backbones to boost their representation ability to detect oriented objects accurately. Experiments on commonly used benchmarks (DOTA and HRSC2016) demonstrate that equipped with our proposed ARC module in the backbone network, the performance of multiple popular oriented object detectors is significantly improved (e.g. +3.03% mAP on Rotated RetinaNet and +4.16% on CFA). Combined with the highly competitive method Oriented R-CNN, the proposed approach achieves state-of-the-art performance on the DOTA dataset with 81.77% mAP. Code is available at https://github.com/LeapLabTHU/ARC.
Yifan Pu, Yiru Wang 0003, Zhuofan Xia, Yizeng Han, Yulin Wang 0002, Weihao Gan, Zidong Wang 0011, Shiji Song, Gao Huang 0001
ICCV8
2023 EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual Backbones
abstract
The superior performance of modern deep networks usually comes with a costly training procedure. This paper presents a new curriculum learning approach for the efficient training of visual backbones (e.g., vision Transformers). Our work is inspired by the inherent learning dynamics of deep networks: we experimentally show that at an earlier training stage, the model mainly learns to recognize some ‘easier-to-learn’ discriminative patterns within each example, e.g., the lower-frequency components of images and the original information before data augmentation. Driven by this phenomenon, we propose a curriculum where the model always leverages all the training data at each epoch, while the curriculum starts with only exposing the ‘easier-to-learn’ patterns of each example, and introduces gradually more difficult patterns. To implement this idea, we 1) introduce a cropping operation in the Fourier spectrum of the inputs, which enables the model to learn from only the lower-frequency components efficiently, 2) demonstrate that exposing the features of original images amounts to adopting weaker data augmentation, and 3) integrate 1) and 2) and design a curriculum learning schedule with a greedy-search algorithm. The resulting approach, EfficientTrain, is simple, general, yet surprisingly effective. As an off-the-shelf method, it reduces the wall-time training cost of a wide variety of popular models (e.g., ResNet, ConvNeXt, DeiT, PVT, Swin, and CSWin) by > 1.5× on ImageNet-1K/22K without sacrificing accuracy. It is also effective for self-supervised learning (e.g., MAE). Code is available at https://github.com/LeapLabTHU/EfficientTrain.
Yulin Wang 0002, Rui Lu 0001, Zhao Zhong, Shiji Song, Gao Huang 0001
ICCV6
2023 Budgeted Training for Vision Transformer
Zhuofan Xia, Xuran Pan, Xuan Jin, Yuan He 0011, Hui Xue 0001, Shiji Song, Gao Huang 0001
ICLR6
2023 Boosting Offline Reinforcement Learning with Action Preference Query
abstract
Training practical agents usually involve offline and online reinforcement learning (RL) to balance the policy's performance and interaction costs. In particular, online fine-tuning has become a commonly used method to correct the erroneous estimates of out-of-distribution data learned in the offline training phase. However, even limited online interactions can be inaccessible or catastrophic for high-stake scenarios like healthcare and autonomous driving. In this work, we introduce an interaction-free training scheme dubbed Offline-with-Action-Preferences (OAP). The main insight is that, compared to online fine-tuning, querying the preferences between pre-collected and learned actions can be equally or even more helpful to the erroneous estimate problem. By adaptively encouraging or suppressing policy constraint according to action preferences, OAP could distinguish overestimation from beneficial policy improvement and thus attains a more accurate evaluation of unseen data. Theoretically, we prove a lower bound of the behavior policy's performance improvement brought by OAP. Moreover, comprehensive experiments on the D4RL benchmark and state-of-the-art algorithms demonstrate that OAP yields higher (29% on average) scores, especially on challenging AntMaze tasks (98% higher).
Qisen Yang, Shenzhi Wang, Matthieu Lin, Shiji Song, Gao Huang 0001
ICML4
2023 Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement Learning
abstract
Offline-to-online reinforcement learning (RL) is a training paradigm that combines pre-training on a pre-collected dataset with fine-tuning in an online environment. However, the incorporation of online fine-tuning can intensify the well-known distributional shift problem. Existing solutions tackle this problem by imposing a policy constraint on the policy improvement objective in both offline and online learning. They typically advocate a single balance between policy improvement and constraints across diverse data collections. This one-size-fits-all manner may not optimally leverage each collected sample due to the significant variation in data quality across different states. To this end, we introduce Family Offline-to-Online RL (FamO2O), a simple yet effective framework that empowers existing algorithms to determine state-adaptive improvement-constraint balances. FamO2O utilizes a universal model to train a family of policies with different improvement/constraint intensities, and a balance model to select a suitable policy for each state. Theoretically, we prove that state-adaptive balances are necessary for achieving a higher policy performance upper bound. Empirically, extensive experiments show that FamO2O offers a statistically significant improvement over various existing methods, achieving state-of-the-art performance on the D4RL benchmark. Codes are available at https://github.com/LeapLabTHU/FamO2O.
Shenzhi Wang, Qisen Yang, Jiawei Gao 0004, Matthieu Lin, Shiji Song, Gao Huang 0001
NeurIPS8
2023 Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL
abstract
The divergence of the Q-value estimation has been a prominent issue offline reinforcement learning (offline RL), where the agent has no access to real dynamics. Traditional beliefs attribute this instability to querying out-of-distribution actions when bootstrapping value targets. Though this issue can be alleviated with policy constraints or conservative Q estimation, a theoretical understanding of the underlying mechanism causing the divergence has been absent. In this work, we aim to thoroughly comprehend this mechanism and attain an improved solution. We first identify a fundamental pattern, \emph{self-excitation}, as the primary cause of Q-value estimation divergence in offline RL. Then, we propose a novel Self-Excite Eigenvalue Measure (SEEM) metric based on Neural Tangent Kernel (NTK) to measure the evolving property of Q-network at training, which provides an intriguing explanation of the emergence of divergence. For the first time, our theory can reliably decide whether the training will diverge at an early stage, and even predict the order of the growth for the estimated Q-value, the model's norm, and the crashing step when an SGD optimizer is used. The experiments demonstrate perfect alignment with this theoretic analysis. Building on our insights, we propose to resolve divergence from a novel perspective, namely improving the model's architecture for better extrapolating behavior. Through extensive empirical studies, we identify LayerNorm as a good solution to effectively avoid divergence without introducing detrimental bias, leading to superior performance. Experimental results prove that it can still work in some most challenging settings, i.e. using only 1$\%$ transitions of the dataset, where all previous methods fail. Moreover, it can be easily plugged into modern offline RL methods and achieve SOTA results on many challenging tasks. We also give unique insights into its effectiveness.
Rui Lu 0001, Bingyi Kang, Shiji Song, Gao Huang 0001
NeurIPS4
2023 Accelerating Column Generation Algorithm Using Machine-Learning-Based Column Elimination
abstract
The column generation (CG) algorithm is widely used in large-scale optimization problems. However, a large amount of columns in the restricted master problem (RMP) makes the computing process very time-consuming. This paper proposes a machine learning based column elimination strategy to accelerate the CG algorithm. Our approach represents the RMP by a bipartite graph and applies a learned Graph Neural Network model to predict redundant columns to be eliminated from the RMP, so as to reduce the time cost of solving the RMP and iterations required for convergence. Our approach is tested on cutting stock problem instances. Compared with the vanilla CG algorithm, the iterations and time required for convergence are reduced by up to 31 % and 48%, respectively. Furthermore, our approach shows great generalization to cutting stock problem instances of different sizes.
Lichang Fang, Haofeng Yuan, Shiji Song
SMC4
2023 MLP-based classification of COVID-19 and skin diseases
Ruize Zhang 0002, Shuli Cheng, Shiji Song
Expert Syst. Appl.4
2023 Synchronization of switched complex dynamical networks with impulses: state-dependent switching approach
Dan Yang 0013, Xiaodi Li 0001, Shiji Song
Neurocomputing3
2023 Glance and Focus Networks for Dynamic Visual Recognition
abstract
Spatial redundancy widely exists in visual recognition tasks, i.e., discriminative features in an image or video frame usually correspond to only a subset of pixels, while the remaining regions are irrelevant to the task at hand. Therefore, static models which process all the pixels with an equal amount of computation result in considerable redundancy in terms of time and space consumption. In this paper, we formulate the image recognition problem as a sequential coarse-to-fine feature learning process, mimicking the human visual system. Specifically, the proposed Glance and Focus Network (GFNet) first extracts a quick global representation of the input image at a low resolution scale, and then strategically attends to a series of salient (small) regions to learn finer features. The sequential process naturally facilitates adaptive inference at test time, as it can be terminated once the model is sufficiently confident about its prediction, avoiding further redundant computation. It is worth noting that the problem of locating discriminant regions in our model is formulated as a reinforcement learning task, thus requiring no additional manual annotations other than classification labels. GFNet is general and flexible as it is compatible with any off-the-shelf backbone models (such as MobileNets, EfficientNets and TSM), which can be conveniently deployed as the feature extractor. Extensive experiments on a variety of image classification and video recognition tasks and with various backbone models demonstrate the remarkable efficiency of our method. For example, it reduces the average latency of the highly efficient MobileNet-V3 on an iPhone XS Max by 1.3x without sacrificing accuracy. Code and pre-trained models are available at https://github.com/blackfeather-wang/GFNet-Pytorch.
Gao Huang 0001, Yulin Wang 0002, Kangchen Lv, Haojun Jiang, Shiji Song
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Fast trajectory planning for Dubins vehicles under cumulative probability of radar detection
Zhuo Li 0011, Keyou You, Jian Sun 0003, Shiji Song
Signal Process.4
2023 The Hot Strip Mill Scheduling Problem With Uncertainty: Robust Optimization Models and Solution Approaches
abstract
In this article, we focus on a biobjective hot strip mill (HSM) scheduling problem arising in the steel industry. Besides the conventional objective regarding penalty costs, we have also considered minimizing the total starting times of rolling operations in order to reduce the energy consumption for slab reheating. The problem is complicated by the inevitable uncertainty in rolling processing times, which means deterministic scheduling models will be ineffective. To obtain robust production schedules with satisfactory performance under all possible conditions, we apply the robust optimization (RO) approach to model and solve the scheduling problem. First, an RO model and an equivalent mixed-integer linear programming model are constructed to describe the HSM scheduling problem with uncertainty. Then, we devise an improved Benders' decomposition algorithm to solve the RO model and obtain exactly optimal solutions. Next, for coping with large-sized instances, a multiobjective particle swarm optimization algorithm with an embedded local search strategy is proposed to handle the biobjective scheduling problem and find the set of Pareto-optimal solutions. Finally, we conduct extensive computational tests to verify the proposed algorithms. Results show that the exact algorithm is effective for relatively small instances and the metaheuristic algorithm can achieve satisfactory solution quality for both small- and large-sized instances of the problem.
Rui Zhang 0039, Shiji Song, Cheng Wu 0002
IEEE Trans. Cybern.2
2023 Input-to-State Stability of Nonlinear Impulsive Systems Subjects to Actuator Saturation and External Disturbance
abstract
This article mainly explores the local input-to-state stability (LISS) property of a class of nonlinear systems via a saturated control strategy, where both the external disturbance and impulsive disturbance being fully considered. In terms of the Lyapunov method and inequality techniques, some sufficient conditions under which the system can be made LISS are proposed, and the elastic constraint relationship among saturated control gain, rate coefficients, external disturbance, and domain of initial value is revealed. Moreover, the optimization design procedures are provided with the hope of obtaining the estimates of admissible external disturbance and domain of initial value as large as possible, where the corresponding saturated control law can be designed by solving LMI -based conditions. In the absence of an external disturbance, the locally exponential stability (LES) property can also be presented with a set of more relaxed conditions. Finally, two examples are presented to reveal the validity of the obtained results.
Xiaodi Li 0001, Shiji Song
IEEE Trans. Cybern.3
2023 Underwater Attentional Generative Adversarial Networks for Image Enhancement
abstract
In this article, to exclusively suppress unuseful underwater noise feature and effectively avoid overenhancement, simultaneously, an underwater attentional generative adversarial network (UAGAN) is innovatively established. Main contributions are as follows: combining dense concatenation with global maximum and average pooling techniques, a cascade dense-channel attention (CDCA) module is devised to adaptively distinguish noise feature and recalibrate channel weight, simultaneously, such that low-contribution feature map can be effectively suppressed; to sufficiently capture long-range dependence between any two nonlocal spatial patches, the position attention (PA) module is created such that the deviation among independent patches can be sufficiently eliminated, thereby avoiding overenhancement; and in conjunction with CDCA and PA modules, the entire UAGAN framework is eventually developed in an end-to-end manner. Comprehensive experiments conducted on underwater image enhancement benchmark (UIEB) and underwater robot professional contest (URPC) datasets demonstrate remarkable effectiveness and superiority of the proposed UAGAN scheme by comparing with typical underwater image enhancement approaches including unsupervised color correction method, image blurriness and light absorption, underwater dark channel prior, underwater generative adversarial network, underwater convolutional neural network, and WaterNet in terms of peak signal-to-noise ratio, underwater color image quality evaluation, underwater image quality measures, etc.
Ning Wang 0002, Tingkai Chen, Xiangjun Kong, Rongfeng Wang, Yongjun Gong, Shiji Song
IEEE Trans. Hum. Mach. Syst.7
2023 Meta-Reinforcement Learning With Dynamic Adaptiveness Distillation
abstract
Deep reinforcement learning is confronted with problems of sampling inefficiency and poor task migration capability. Meta-reinforcement learning (meta-RL) enables meta-learners to utilize the task-solving skills trained on similar tasks and quickly adapt to new tasks. However, meta-RL methods lack enough queries toward the relationship between task-agnostic exploitation of data and task-related knowledge introduced by latent context, limiting their effectiveness and generalization ability. In this article, we develop an algorithm for off-policy meta-RL that can provide the meta-learners with self-oriented cognition toward how they adapt to the family of tasks. In our approach, we perform dynamic task-adaptiveness distillation to describe how the meta-learners adjust the exploration strategy in the meta-training process. Our approach also enables the meta-learners to balance the influence of task-agnostic self-oriented adaption and task-related information through latent context reorganization. In our experiments, our method achieves 10%-20% higher asymptotic reward than probabilistic embeddings for actor-critic RL (PEARL).
Hangkai Hu, Gao Huang 0001, Xiang Li 0009, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.4
2023 Exploration With Task Information for Meta Reinforcement Learning
abstract
Meta reinforcement learning (meta-RL) is a promising technique for fast task adaptation by leveraging prior knowledge from previous tasks. Recently, context-based meta-RL has been proposed to improve data efficiency by applying a principled framework, dividing the learning procedure into task inference and task execution. However, the task information is not adequately leveraged in this approach, thus leading to inefficient exploration. To address this problem, we propose a novel context-based meta-RL framework with an improved exploration mechanism. For the existing exploration and execution problem in context-based meta-RL, we propose a novel objective that employs two exploration terms to encourage better exploration in action and task embedding space, respectively. The first term pushes for improving the diversity of task inference, while the second term, named action information, works as sharing or hiding task information in different exploration stages. We divide the meta-training procedure into task-independent exploration and task-relevant exploration stages according to the utilization of action information. By decoupling task inference and task execution and proposing the respective optimization objectives in the two exploration stages, we can efficiently learn policy and task inference networks. We compare our algorithm with several popular meta-RL methods on MuJoco benchmarks with both dense and sparse reward settings. The empirical results show that our method significantly outperforms baselines on the benchmarks in terms of sample efficiency and task performance.
Peng Jiang 0011, Shiji Song, Gao Huang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Global Model Learning for Large Deformation Control of Elastic Deformable Linear Objects: An Efficient and Adaptive Approach
abstract
The robotic manipulation of deformable linear objects (DLOs) has broad application prospects in many fields. However, a key issue is to obtain the exact deformation models (i.e., how robot motion affects DLO deformation), which are hard to theoretically calculate and vary among different DLOs. Thus, the shape control of DLOs is challenging, especially for large deformation control that requires global and more accurate models. In this article, we propose a coupled offline and online data-driven method for efficiently learning a global deformation model, allowing for both accurate modeling through offline learning and further updating for new DLOs via online adaptation. Specifically, the model approximated by a neural network is first trained offline on random data, then seamlessly migrated to the online phase, and further updated online during actual manipulation. Several strategies are introduced to improve the model's efficiency and generalization ability. We propose a convex-optimization-based controller and analyze the system's stability using the Lyapunov method. Detailed simulations and real-world experiments demonstrate that our method can efficiently and precisely estimate the deformation model and achieve the large deformation control of untrained DLOs in 2-D and 3-D dual-arm manipulation tasks better than the existing methods. It accomplishes all 24 tasks with different desired shapes on different DLOs in the real world, using only simulation data for the offline learning.
Mingrui Yu 0001, Kangchen Lv, Hanzhong Zhong, Shiji Song, Xiang Li 0009
IEEE Trans. Robotics4
2023 Prescribed-Time Stabilization of Nonlinear Systems via Impulsive Regulation
abstract
This article studies the problem of prescribed-time stabilization for nonlinear systems, in which impulses are cautiously regulated not only to stabilize the system in a finite-time sense but also to adjust the settling time to fit the prescribed terminal time. By fetching and utilizing positive effect of impulses, constrains on continuous flows are effectively relaxed for prescribed-time stabilization of nonlinear systems. Meanwhile, different from traditional prescribed-time approaches, where continuous flows of the system are generally required to be globally stable, it shows in this article that even systems involving unstable flows can be prescribed-time stabilized via the proposed impulsive regulation, but as a tradeoff, a constrained initial condition is needed for controller design. Thus, the main functions of impulsive regulation in this article are balancing unstable dynamics of the system, and regulating the settling time to fit the prescribed terminal time. As an application, the proposed impulsive regulation scheme is exploited in prescribed-time stabilization of affine dynamical systems involving uncertain input noises. Two examples, including the one focusing on the spin stabilization of spacecrafts, are given to verify the results.
Xiaodi Li 0001, Shiji Song
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding
abstract
Visual grounding, i.e., localizing objects in images ac-cording to natural language queries, is an important topic in visual language understanding. The most effective approaches for this task are based on deep learning, which generally require expensive manually labeled image-query or patch-query pairs. To eliminate the heavy depen-dence on human annotations, we present a novel method, named Pseudo-Q, to automatically generate pseudo language queries for supervised training. Our method lever-ages an off-the-shelf object detector to identify visual ob-jects from unlabeled images, and then language queries for these objects are obtained in an unsupervised fashion with a pseudo-query generation module. Then, we design a task-related query prompt module to specifically tailor generated pseudo language queries for visual grounding tasks. Further, in order to fully capture the contextual re-lationships between images and language queries, we de-velop a visual-language model equipped with multi-level cross-modality attention mechanism. Extensive experimen-tal results demonstrate that our method has two notable benefits: (1) it can reduce human annotation costs signifi-cantly, e.g., 31% on Ref Coco [65] without degrading orig-inal model's performance under the fully supervised set-ting, and (2) without bells and whistles, it achieves supe-rior or comparable performance compared to state-of-the-art weakly-supervised visual grounding methods on all the five datasets we have experimented. Code is available at https://github.com/LeapLabTHU/Pseudo-Q.
Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, Gao Huang 0001
CVPR4
2022 On the Integration of Self-Attention and Convolution
abstract
Convolution and self-attention are two powerful techniques for representation learning, and they are usually considered as two peer approaches that are distinct from each other. In this paper, we show that there exists a strong underlying relation between them, in the sense that the bulk of computations of these two paradigms are in fact done with the same operation. Specifically, we first show that a traditional convolution with kernel size k × k can be decomposed into k2individual 1 × 1 convolutions, followed by shift and summation operations. Then, we interpret the projections of queries, keys, and values in self-attention module as multiple 1 × 1 convolutions, followed by the computation of attention weights and aggregation of the values. Therefore, the first stage of both two modules comprises the similar operation. More importantly, the first stage contributes a dominant computation complexity (square of the channel size) comparing to the second stage. This observation naturally leads to an elegant integration of these two seemingly distinct paradigms, i.e., a mixed model that enjoys the benefit of both self-Attention and Convolution (ACmix), while having minimum compu-tational overhead compared to the pure convolution or self-attention counterpart. Extensive experiments show that our model achieves consistently improved results over com-petitive baselines on image recognition and downstream tasks. Code and pre-trained models will be released at https://github.com/LeapLabTHU/ACmix and https://gitee.com/mindspore/models.
Xuran Pan, Chunjiang Ge, Rui Lu 0001, Shiji Song, Guanfu Chen, Zeyi Huang, Gao Huang 0001
CVPR4
2022 Exploring the Equivalence of Siamese Self-Supervised Learning via A Unified Gradient Framework
abstract
Self-supervised learning has shown its great potential to extract powerful visual representations without human annotations. Various works are proposed to deal with self-supervised learning from different perspectives: (1) contrastive learning methods (e.g., MoCo, SimCLR) utilize both positive and negative samples to guide the training direction; (2) asymmetric network methods (e.g., BYOL, SimSiam) get rid of negative samples via the introduction of a predictor network and the stop-gradient operation; (3) feature decorrelation methods (e.g., Barlow Twins, VICReg) instead aim to reduce the redundancy between feature dimensions. These methods appear to be quite different in the designed loss functions from various motivations. The final accuracy numbers also vary, where different networks and tricks are utilized in different works. In this work, we demonstrate that these methods can be unified into the same form. Instead of comparing their loss functions, we derive a unified formula through gradient analysis. Furthermore, we conduct fair and detailed experiments to compare their performances. It turns out that there is little gap between these methods, and the use of momentum encoder is the key factor to boost performance. From this unified framework, we propose UniGrad, a simple but effective gradient form for self-supervised learning. It does not require a memory bank or a predictor network, but can still achieve state-of-the-art performance and easily adopt other training strategies. Extensive experiments on linear evaluation and many downstream tasks also show its effectiveness. Code shall be released.
Chenxin Tao, Xizhou Zhu, Jiahua Dong 0002, Shiji Song, Gao Huang 0001, Jifeng Dai
CVPR5
2022 Vision Transformer with Deformable Attention
abstract
Transformers have recently shown superior performances on various vision tasks. The large, sometimes even global, receptive field endows Transformer models with higher representation power over their CNN counterparts. Nevertheless, simply enlarging receptive field also gives rise to several concerns. On the one hand, using dense attention e.g., in ViT, leads to excessive memory and computational cost, and features can be influenced by irrelevant parts which are beyond the region of interests. On the other hand, the sparse attention adopted in PVT or Swin Transformer is data agnostic and may limit the ability to model long range relations. To mitigate these issues, we propose a novel deformable selfattention module, where the positions of key and value pairs in selfattention are selected in a data-dependent way. This flexible scheme enables the self-attention module to focus on relevant re-gions and capture more informative features. On this basis, we present Deformable Attention Transformer, a general backbone model with deformable attention for both image classification and dense prediction tasks. Extensive experi-ments show that our models achieve consistently improved results on comprehensive benchmarks. Code is available at https://github.com/LeapLabTHU/DAT.
Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, Gao Huang 0001
CVPR3
2022 Learning to Weight Samples for Dynamic Early-Exiting Networks
Yizeng Han, Yifan Pu, Zihang Lai, Chaofei Wang, Shiji Song, Junfen Cao, Chao Deng 0002, Gao Huang 0001
ECCV (11)5
2022 ActiveNeRF: Learning Where to See with Uncertainty Estimation
Xuran Pan, Zihang Lai, Shiji Song, Gao Huang 0001
ECCV (33)3
2022 AdaFocusV3: On Unified Spatial-Temporal Dynamic Video Recognition
Yulin Wang 0002, Xinhong Xu, Ali Hassani 0001, Victor Kulikov, Nikita Orlov, Shiji Song, Humphrey Shi, Gao Huang 0001
ECCV (4)7
2022 Learn From the Past: Experience Ensemble Knowledge Distillation
abstract
Traditional knowledge distillation transfers "dark knowledge" of a pre-trained teacher network to a student network, and ignores the knowledge in the training process of the teacher, which we call teacher’s experience. However, in realistic educational scenarios, learning experience is often more important than learning results. In this work, we propose a novel knowledge distillation method by integrating the teacher’s experience for knowledge transfer, named experience ensemble knowledge distillation (EEKD). We save a moderate number of intermediate models from the training process of the teacher model uniformly, and then integrate the knowledge of these intermediate models by ensemble technique. A self-attention module is used to adaptively assign weights to different intermediate models in the process of knowledge transfer. Three principles of constructing EEKD on the quality, weights and number of intermediate models are explored. A surprising conclusion is found that strong ensemble teachers do not necessarily produce strong students. The experimental results on CIFAR-100 and ImageNet show that EEKD outperforms the mainstream knowledge distillation methods and achieves the state-of-the-art. In particular, EEKD even surpasses the standard ensemble distillation on the premise of saving training cost.
Chaofei Wang, Shiji Song, Gao Huang 0001
ICPR3
2022 Latency-aware Spatial-wise Dynamic Networks
abstract
Spatial-wise dynamic convolution has become a promising approach to improving the inference efficiency of deep networks. By allocating more computation to the most informative pixels, such an adaptive inference paradigm reduces the spatial redundancy in image features and saves a considerable amount of unnecessary computation. However, the theoretical efficiency achieved by previous methods can hardly translate into a realistic speedup, especially on the multi-core processors (e.g. GPUs). The key challenge is that the existing literature has only focused on designing algorithms with minimal computation, ignoring the fact that the practical latency can also be influenced by scheduling strategies and hardware properties. To bridge the gap between theoretical computation and practical efficiency, we propose a latency-aware spatial-wise dynamic network (LASNet), which performs coarse-grained spatially adaptive inference under the guidance of a novel latency prediction model. The latency prediction model can efficiently estimate the inference latency of dynamic networks by simultaneously considering algorithms, scheduling strategies, and hardware properties. We use the latency predictor to guide both the algorithm design and the scheduling optimization on various hardware platforms. Experiments on image classification, object detection and instance segmentation demonstrate that the proposed framework significantly improves the practical inference efficiency of deep networks. For example, the average latency of a ResNet-101 on the ImageNet validation set could be reduced by 36% and 46% on a server GPU (Nvidia Tesla-V100) and an edge device (Nvidia Jetson TX2 GPU) respectively without sacrificing the accuracy. Code is available at https://github.com/LeapLabTHU/LASNet.
Yizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue, Shiji Song, Guangyu Sun 0003, Gao Huang 0001
NeurIPS5
2022 Contrastive Language-Image Pre-Training with Knowledge Graphs
abstract
Recent years have witnessed the fast development of large-scale pre-training frameworks that can extract multi-modal representations in a unified form and achieve promising performances when transferred to downstream tasks. Nevertheless, existing approaches mainly focus on pre-training with simple image-text pairs, while neglecting the semantic connections between concepts from different modalities. In this paper, we propose a knowledge-based pre-training framework, dubbed Knowledge-CLIP, which injects semantic information into the widely used CLIP model. Through introducing knowledge-based objectives in the pre-training process and utilizing different types of knowledge graphs as training data, our model can semantically align the representations in vision and language with higher quality, and enhance the reasoning ability across scenarios and modalities. Extensive experiments on various vision-language downstream tasks demonstrate the effectiveness of Knowledge-CLIP compared with the original CLIP and competitive baselines.
Xuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song, Gao Huang 0001
NeurIPS4
2022 Efficient Knowledge Distillation from Model Checkpoints
abstract
Knowledge distillation is an effective approach to learn compact models (students) with the supervision of large and strong models (teachers). As empirically there exists a strong correlation between the performance of teacher and student models, it is commonly believed that a high performing teacher is preferred. Consequently, practitioners tend to use a well trained network or an ensemble of them as the teacher. In this paper, we observe that an intermediate model, i.e., a checkpoint in the middle of the training procedure, often serves as a better teacher compared to the fully converged model, although the former has much lower accuracy. More surprisingly, a weak snapshot ensemble of several intermediate models from a same training trajectory can outperform a strong ensemble of independently trained and fully converged models, when they are used as teachers. We show that this phenomenon can be partially explained by the information bottleneck principle: the feature representations of intermediate models can have higher mutual information regarding the input, and thus contain more ``dark knowledge'' for effective distillation. We further propose an optimal intermediate teacher selection algorithm based on maximizing the total task-related mutual information. Experiments verify its effectiveness and applicability. Our code is available at https://github.com/LeapLabTHU/CheckpointKD.
Chaofei Wang, Qisen Yang, Rui Huang 0012, Shiji Song, Gao Huang 0001
NeurIPS4
2022 The Neural-Prediction based Acceleration Algorithm of Column Generation for Graph-Based Set Covering Problems
abstract
Set covering problem is an important class of combinatorial optimization problems, which has been widely applied and studied in many fields. In this paper, we propose an improved column generation algorithm with neural prediction (CG-P) for solving graph-based set covering problems. We leverage a graph neural network based neural prediction model to predict the probability to be included in the final solution for each edge. Our CG-P algorithm constructs a reduced graph that only contains the edges with higher predicted probability, and this graph reduction process significantly speeds up the solution process. We evaluate the CG-P algorithm on railway crew scheduling problems and it outperforms the baseline column generation algorithm. We provide two solution modes for our CG-P algorithm. In the optimal mode, we can obtain a solution with an optimality guarantee while reducing the time cost to 63.12%. In the fast mode, we can obtain a sub-optimal solution with a 7.62% optimality gap in only 2.91% computation time.
Haofeng Yuan, Peng Jiang 0011, Shiji Song
SMC3
2022 PLAM: A plug-in module for flexible graph attention learning
Xuran Pan, Shiji Song, Yiming Chen 0005, Gao Huang 0001
Neurocomputing2
2022 TC3KD: Knowledge distillation via teacher-student cooperative curriculum customization
Chaofei Wang, Gao Huang 0001, Shiji Song
Neurocomputing5
2022 Finite-time stability of state-dependent delayed systems and application to coupled neural networks
Xiaodi Li 0001, Shiji Song
Neural Networks3
2022 Dynamic Neural Networks: A Survey
abstract
Dynamic neural network is an emerging research topic in deep learning. Compared to static models which have fixed computational graphs and parameters at the inference stage, dynamic networks can adapt their structures or parameters to different inputs, leading to notable advantages in terms of accuracy, computational efficiency, adaptiveness, etc. In this survey, we comprehensively review this rapidly developing area by dividing dynamic networks into three main categories: 1) sample-wise dynamic models that process each sample with data-dependent architectures or parameters; 2) spatial-wise dynamic networks that conduct adaptive computation with respect to different spatial locations of image data; and 3) temporal-wise dynamic models that perform adaptive inference along the temporal dimension for sequential data such as videos and texts. The important research problems of dynamic networks, e.g., architecture design, decision making scheme, optimization technique and applications, are reviewed systematically. Finally, we discuss the open problems in this field together with interesting future research directions.
Yizeng Han, Gao Huang 0001, Shiji Song, Le Yang 0007, Yulin Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Temporal-Spatial Causal Interpretations for Vision-Based Reinforcement Learning
abstract
Deep reinforcement learning (RL) agents are becoming increasingly proficient in a range of complex control tasks. However, the agent's behavior is usually difficult to interpret due to the introduction of black-box function, making it difficult to acquire the trust of users. Although there have been some interesting interpretation methods for vision-based RL, most of them cannot uncover temporal causal information, raising questions about their reliability. To address this problem, we present a temporal-spatial causal interpretation (TSCI) model to understand the agent's long-term behavior, which is essential for sequential decision-making. TSCI model builds on the formulation of temporal causality, which reflects the temporal causal relations between sequential observations and decisions of RL agent. Then a separate causal discovery network is employed to identify temporal-spatial causal features, which are constrained to satisfy the temporal causality. TSCI model is applicable to recurrent agents and can be used to discover causal features with high efficiency once trained. The empirical results show that TSCI model can produce high-resolution and sharp attention masks to highlight task-relevant temporal-spatial information that constitutes most evidence about how vision-based RL agents make sequential decisions. In addition, we further demonstrate that our method is able to provide valuable causal interpretations for vision-based RL agents from the temporal perspective.
Wenjie Shi, Gao Huang 0001, Shiji Song, Cheng Wu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Self-Supervised Discovering of Interpretable Features for Reinforcement Learning
abstract
Deep reinforcement learning (RL) has recently led to many breakthroughs on a range of complex control tasks. However, the agent's decision-making process is generally not transparent. The lack of interpretability hinders the applicability of RL in safety-critical scenarios. While several methods have attempted to interpret vision-based RL, most come without detailed explanation for the agent's behavior. In this paper, we propose a self-supervised interpretable framework, which can discover interpretable features to enable easy understanding of RL agents even for non-experts. Specifically, a self-supervised interpretable network (SSINet) is employed to produce fine-grained attention masks for highlighting task-relevant information, which constitutes most evidence for the agent's decisions. We verify and evaluate our method on several Atari 2600 games as well as Duckietown, which is a challenging self-driving car simulator environment. The results show that our method renders empirical evidences about how the agent makes decisions and why the agent performs well or badly, especially when transferred to novel scenes. Overall, our method provides valuable insight into the internal decision-making process of vision-based RL. In addition, our method does not use any external labelled data, and thus demonstrates the possibility to learn high-quality mask through a self-supervised manner, which may shed light on new paradigms for label-free vision learning such as self-supervised segmentation and detection.
Wenjie Shi, Gao Huang 0001, Shiji Song, Tingyu Lin 0001, Cheng Wu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Regularizing Deep Networks With Semantic Data Augmentation
abstract
Data augmentation is widely known as a simple yet surprisingly effective technique for regularizing deep networks. Conventional data augmentation schemes, e.g., flipping, translation or rotation, are low-level, data-independent and class-agnostic operations, leading to limited diversity for augmented samples. To this end, we propose a novel semantic data augmentation algorithm to complement traditional approaches. The proposed method is inspired by the intriguing property that deep networks are effective in learning linearized features, i.e., certain directions in the deep feature space correspond to meaningful semantic transformations, e.g., changing the background or view angle of an object. Based on this observation, translating training samples along many such directions in the feature space can effectively augment the dataset for more diversity. To implement this idea, we first introduce a sampling based method to obtain semantically meaningful directions efficiently. Then, an upper bound of the expected cross-entropy (CE) loss on the augmented training set is derived by assuming the number of augmented samples goes to infinity, yielding a highly efficient algorithm. In fact, we show that the proposed implicit semantic data augmentation (ISDA) algorithm amounts to minimizing a novel robust CE loss, which adds minimal extra computational cost to a normal training procedure. In addition to supervised learning, ISDA can be applied to semi-supervised learning tasks under the consistency regularization framework, where ISDA amounts to minimizing the upper bound of the expected KL-divergence between the augmented features and the original features. Although being simple, ISDA consistently improves the generalization performance of popular deep models (e.g., ResNets and DenseNets) on a variety of datasets, i.e., CIFAR-10, CIFAR-100, SVHN, ImageNet, and Cityscapes. Code for reproducing our results is available at https://github.com/blackfeather-wang/ISDA-for-Deep-Networks.
Yulin Wang 0002, Gao Huang 0001, Shiji Song, Xuran Pan, Yitong Xia, Cheng Wu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Second-Order Conic Programming Approach for Wasserstein Distributionally Robust Two-Stage Linear Programs
abstract
This article proposes a second-order conic programming (SOCP) approach to solve distributionally robust two-stage linear programs over 1-Wasserstein balls. We start from the case with distribution uncertainty only in the objective function and then explore the case with distribution uncertainty only in constraints. The former program is exactly reformulated as a tractable SOCP problem, whereas the latter one is proved to be generally NP-hard as it involves a norm maximization problem over a polyhedron. However, it reduces to an SOCP problem if the extreme points of the polyhedron are given as a prior. This motivates the design of a constraint generation algorithm with provable convergence to approximately solve the NP-hard problem. Moreover, the least favorable distribution achieving the worst case cost is given as an “empirical” distribution by simply perturbing each original sample for both cases. Finally, experiments illustrate the advantages of the proposed model in terms of the out-of-sample performance and computational complexity.Note to Practitioners—The two-stage program with distribution uncertainty is an important decision problem in broad applications, e.g., two-stage schedule problems, facility location problems, and recourse allocation problems. To deal with the uncertainty, this work proposes a novel data-driven model over the 1-Wasserstein ball and develops an efficient second-order conic programming (SOCP)-based solution approach, where the sample data set can be easily exploited to reduce the distribution uncertainty. The good out-of-sample performance and computational complexity of the proposed model are validated by the experiments on the two-stage portfolio programs and material order programs.
Zhuolin Wang, Keyou You, Shiji Song
IEEE Trans Autom. Sci. Eng.3
2022 A Hybrid Artificial Immune-Simulated Annealing Algorithm for Multiroute Job Shop Scheduling Problem With Continuous Limited Output Buffers
abstract
In this article, we study the multiroute job shop scheduling problem with continuous-limited output buffers (MRJSP-CLOBs). In contrast to the standard job shop scheduling problem (JSP), continuous-limited output buffers render the commonly used graph-based approaches inapplicable, and the multiroute issue further increases computational complexity. To this end, we formulate MRJSP-CLOB as a mixed-integer linear program (MILP), which is typically NP-hard. Then, we extend the critical block in the JSP by utilizing the no-time-gap relationship and design a new neighborhood structure. Furthermore, we propose a hybrid artificial immune-simulated annealing algorithm (AIA-SA) by sharing iterations and integrating a random infeasible solution repairing algorithm with a new SA acceptance rule, which enables individuals to share information and increases the robustness of the corresponding SA parameters. Finally, the AIA-SA is compared with CPLEX and state-of-the-art algorithms on MRJSP-CLOB with different sizes. Experiments for large-sized instances demonstrate that our algorithm requires less than 3% computing time of the CPLEX, while being faster and more accurate than the other algorithms.
Shiji Song, Shengsheng Niu, Rui Zhang 0039
IEEE Trans. Cybern.2
2022 Attention-Based Meta-Reinforcement Learning for Tracking Control of AUV With Time-Varying Dynamics
abstract
Reinforcement learning (RL) is a promising technique for designing a model-free controller by interacting with the environment. Several researchers have applied RL to autonomous underwater vehicles (AUVs) for motion control, such as trajectory tracking. However, the existing RL-based controller usually assumes that the unknown AUV dynamics keep invariant during the operation period, limiting its further application in the complex underwater environment. In this article, a novel meta-RL-based control scheme is proposed for trajectory tracking control of AUV in the presence of unknown and time-varying dynamics. To this end, we divide the tracking task for AUV with time-varying dynamics into multiple specific tasks with fixed time-varying dynamics, to which we apply meta-RL for training to distill the general control policy. The obtained control policy can transfer to the testing phase with high adaptability. Inspired by the line-of-sight (LOS) tracking rule, we formulate each specific task as a Markov decision process (MDP) with a well-designed state and reward function. Furthermore, a novel policy network with an attention module is proposed to extract the hidden information of AUV dynamics. The simulation environment with time-varying dynamics is established, and the simulation results reveal the effectiveness of our proposed method.
Peng Jiang 0011, Shiji Song, Gao Huang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Finite-Time Synchronization for Delayed Complex Dynamical Networks With Synchronizing or Desynchronizing Impulses
abstract
In this article, the finite-time synchronization problem of delayed complex dynamical networks (CDNs) with impulses is studied, where two types of impulses, namely, synchronizing impulses and desynchronizing impulses, are fully considered, respectively. Since the existence of impulses makes the discontinuity of the states, which means that the classical result for finite-time stability is inapplicable in such a case, the key challenge is how to guarantee the finite-time stability and estimate the settling time in impulse sense. We apply impulsive control theory and finite-time stability theory to CDNs and establish some sufficient conditions for finite-time synchronization, where two kinds of memory controllers are designed for synchronizing impulses and desynchronizing impulses, respectively. Moreover, the upper bounds for settling time of synchronization, which depends on the impulse sequences, are effectively estimated. It shows that the synchronizing impulses can shorten the settling time of synchronization; conversely, the desynchronizing impulses can delay it. Finally, the theoretical analysis is verified by two simulation examples.
Dan Yang 0013, Xiaodi Li 0001, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.3
2022 A Distributionally Robust Scheduling Approach for Uncertain Steelmaking and Continuous Casting Processes
abstract
This article presents a new model to handle the cast break problem caused by small daily disruptions in the processing time of the steelmaking and continuous casting (SCC) production process. In this model, the exact distribution of the uncertain parameters is unknown, and support set, mean, and covariance information is used to describe the uncertain processing time. The problem aims to determine the assignments, sequences, and time points of the charges to be processed on corresponding machines. The main goal is to minimize the expected value of the production objective while reducing the number of cast break occurrences. The problem is solved in two steps. First, a subproblem is developed by fixing the sequences and the assignments of the charges. This subproblem is formulated as a distributionally robust chance-constrained (DRCC) model, in which the constraints are established with certain probabilities even when the uncertain processing times are in their worst cases. A dual approximation method is proposed to convert the model into a semidefinite programming problem so that it can be solved by standard solvers. Additionally, a linear programming approximation method is used to accelerate the solving procedure. A Tabu search algorithm incorporated with a speed-up strategy is also designed to determine the assignments and sequences of the charges. Both simulated data generated from different distributions and actual production data are used to test the efficacy of our model. Results of the numerical experiments show that the schedule obtained from the DRCC model is more robust, i.e., it causes fewer cast breaks than the nominal schedule obtained from a deterministic model.
Shengsheng Niu, Shiji Song, Raymond Chiong
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Finite-Time Stabilization of Switched Systems Under Mode-Dependent Event-Triggered Impulsive Control
abstract
This article studies the event-triggered finite-time stabilization problem of nonlinear switched systems in which the time derivatives of Lyapunov functions of modes are indefinite. Based on a mode-dependent average dwell time constraint and event-triggered impulsive control (ETIC) strategy, a new type of the mode-dependent event-triggered mechanism (MDETM) which can efficiently avoid Zeno behavior for switched systems is presented. Some Lyapunov-based criteria for finite-time stability (FTS) and finite-time contractive stability (FTCS) are obtained, respectively, where a relationship between the prescribed bound and event-triggered mechanism is established. Then, we apply these proposed ETIC strategies to impulsive switched systems and design a class of LMI-based MDETMs to guarantee the FTS/FTCS property. Finally, the effectiveness of presented ETIC strategies is illustrated by two examples.
Taixiang Zhang, Xiaodi Li 0001, Shiji Song
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Smart Train Operation Algorithms Based on Expert Knowledge and Reinforcement Learning
abstract
During decades, the automatic train operation (ATO) system has been gradually adopted in many subway systems for its low-cost and intelligence. This article proposes two smart train operation (STO) algorithms by integrating the expert knowledge with reinforcement learning algorithms. Compared with previous works, the proposed algorithms can realize the control of continuous action for the subway system and optimize multiple critical objectives without using an offline speed profile. First, through learning historical data of experienced subway drivers, we extract the expert knowledge rules and build inference methods to guarantee the riding comfort, the punctuality, and the safety of the subway system. Then we develop two algorithms for optimizing the energy efficiency of train operation. One is the STO algorithm based on deep deterministic policy gradient named (STOD) and the other is the STO algorithm based on normalized advantage function (STON). Finally, we verify the performance of proposed algorithms via some numerical simulations with the real field data from the Yizhuang Line of the Beijing Subway and illustrate that the developed STO algorithm are better than expert manual driving and existing ATO algorithms in terms of energy efficiency. Moreover, STOD and STON can adapt to different trip times and different resistance conditions.
Kaichen Zhou, Shiji Song, Anke Xue, Keyou You, Hui Wu 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2021 3D Object Detection With Pointformer
abstract
Feature learning for 3D object detection from point clouds is very challenging due to the irregularity of 3D point cloud data. In this paper, we propose Pointformer, a Transformer backbone designed for 3D point clouds to learn features effectively. Specifically, a Local Transformer module is employed to model interactions among points in a local region, which learns context-dependent region features at an object level. A Global Transformer is designed to learn context-aware representations at the scene level. To further capture the dependencies among multi-scale representations, we propose Local-Global Transformer to integrate local features with global features from higher resolution. In addition, we introduce an efficient coordinate refinement module to shift down-sampled points closer to object centroids, which improves object proposal generation. We use Pointformer as the backbone for state-of-the-art object detection models and demonstrate significant improvements over original models on both indoor and outdoor datasets.
Xuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li, Gao Huang 0001
CVPR3
2021 CondenseNet V2: Sparse Feature Reactivation for Deep Networks
abstract
Reusing features in deep networks through dense connectivity is an effective way to achieve high computational efficiency. The recent proposed CondenseNet [14] has shown that this mechanism can be further improved if redundant features are removed. In this paper, we propose an alternative approach named sparse feature reactivation (SFR), aiming at actively increasing the utility of features for reusing. In the proposed network, named CondenseNetV2, each layer can simultaneously learn to 1) selectively reuse a set of most important features from preceding layers; and 2) actively update a set of preceding features to increase their utility for later layers. Our experiments show that the proposed models achieve promising performance on image classification (ImageNet and CIFAR) and object detection (MS COCO) in terms of both theoretical efficiency and practical speed.
Le Yang 0007, Haojun Jiang, Ruojin Cai, Yulin Wang 0002, Shiji Song, Gao Huang 0001, Qi Tian 0001
CVPR5
2021 Adaptive Focus for Efficient Video Recognition
abstract
In this paper, we explore the spatial redundancy in video recognition with the aim to improve the computational efficiency. It is observed that the most informative region in each frame of a video is usually a small image patch, which shifts smoothly across frames. Therefore, we model the patch localization problem as a sequential decision task, and propose a reinforcement learning based approach for efficient spatially adaptive video recognition (AdaFocus). In specific, a light-weighted ConvNet is first adopted to quickly process the full video sequence, whose features are used by a recurrent policy network to localize the most task-relevant regions. Then the selected patches are inferred by a high-capacity network for the final prediction. During offline inference, once the informative patch sequence has been generated, the bulk of computation can be done in parallel, and is efficient on modern GPU devices. In addition, we demonstrate that the proposed method can be easily extended by further considering the temporal redundancy, e.g., dynamically skipping less valuable frames. Extensive experiments on five benchmark datasets, i.e., ActivityNet, FCVID, MiniKinetics, Something-Something V1&V2, demonstrate that our method is significantly more efficient than the competitive baselines. Code is available at https://github.com/blackfeather-wang/AdaFocus.
Yulin Wang 0002, Zhaoxi Chen 0007, Haojun Jiang, Shiji Song, Yizeng Han, Gao Huang 0001
ICCV4
2021 Towards Learning Spatially Discriminative Feature Representations
abstract
The backbone of traditional CNN classifier is generally considered as a feature extractor, followed by a linear layer which performs the classification. We propose a novel loss function, termed as CAM-loss, to constrain the embedded feature maps with the class activation maps (CAMs) which indicate the spatially discriminative regions of an image for particular categories. CAM-loss drives the backbone to express the features of target category and suppress the features of non-target categories or background, so as to obtain more discriminative feature representations. It can be simply applied in any CNN architecture with neglectable additional parameters and calculations. Experimental results show that CAM-loss is applicable to a variety of network structures and can be combined with mainstream regularization methods to improve the performance of image classification. The strong generalization ability of CAMloss is validated in the transfer learning and few shot learning tasks. Based on CAM-loss, we also propose a novel CAAM-CAM matching knowledge distillation method. This method directly uses the CAM generated by the teacher network to supervise the CAAM generated by the student network, which effectively improves the accuracy and convergence rate of the student network.
Chaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang, Shiji Song, Gao Huang 0001
ICCV5
2021 Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
Yulin Wang 0002, Zanlin Ni, Shiji Song, Le Yang 0007, Gao Huang 0001
ICLR3
2021 Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition
abstract
Vision Transformers (ViT) have achieved remarkable success in large-scale image recognition. They split every 2D image into a fixed number of patches, each of which is treated as a token. Generally, representing an image with more tokens would lead to higher prediction accuracy, while it also results in drastically increased computational cost. To achieve a decent trade-off between accuracy and speed, the number of tokens is empirically set to 16x16 or 14x14. In this paper, we argue that every image has its own characteristics, and ideally the token number should be conditioned on each individual input. In fact, we have observed that there exist a considerable number of “easy” images which can be accurately predicted with a mere number of 4x4 tokens, while only a small fraction of “hard” ones need a finer representation. Inspired by this phenomenon, we propose a Dynamic Transformer to automatically configure a proper number of tokens for each input image. This is achieved by cascading multiple Transformers with increasing numbers of tokens, which are sequentially activated in an adaptive fashion at test time, i.e., the inference is terminated once a sufficiently confident prediction is produced. We further design efficient feature reuse and relationship reuse mechanisms across different components of the Dynamic Transformer to reduce redundant computations. Extensive empirical results on ImageNet, CIFAR-10, and CIFAR-100 demonstrate that our method significantly outperforms the competitive baselines in terms of both theoretical computational efficiency and practical inference speed. Code and pre-trained models (based on PyTorch and MindSpore) are available at https://github.com/blackfeather-wang/Dynamic-Vision-Transformer and https://github.com/blackfeather-wang/Dynamic-Vision-Transformer-MindSpore.
Yulin Wang 0002, Rui Huang 0012, Shiji Song, Zeyi Huang, Gao Huang 0001
NeurIPS3
2021 Large scale air pollution prediction with deep convolutional networks
Gao Huang 0001, Chunjiang Ge, Tianyu Xiong, Shiji Song, Le Yang 0007, Baoxian Liu, Wenjun Yin, Cheng Wu 0002
Sci. China Inf. Sci.4
2021 Exponential synchronization of delayed neural networks involving unmeasurable neuron states via impulsive observer and impulsive control
Yuhan Wang 0013, Xiaodi Li 0001, Shiji Song
Neurocomputing3
2021 Fine-grained few shot learning with foreground object transformation
Chaofei Wang, Shiji Song, Qisen Yang, Xiang Li 0009, Gao Huang 0001
Neurocomputing2
2021 Fusion layer attention for image-text matching
Depeng Wang, Shiji Song, Gao Huang 0001, Shuli Cheng, Naixiang Ao, Anyu Du
Neurocomputing3
2021 Discriminative Dimension Reduction via Maximin Separation Probability Analysis
abstract
In this paper, we propose a novel discriminative dimension reduction (DR) method, maximin separation probability analysis (MSPA), which maximizes the minimum separation probability of all classes in the reduced low-dimensional subspace. Separation probability is a novel class separability measure, which gives a lower bound of the generalization accuracy for a learned linear classifier in a binary classification problem. The proposed MSPA duly considers the separation of all class pairs in multiclass linear discriminant analysis (LDA) and thus improves the subsequent classification performance. DR via MSPA leads to a nonconvex optimization problem. We develop an algorithm to solve the problem and the global optimal solution can be found by converting the original problem into a series of second-order cone programming problems. A low-computational cost extension and a non-LDA with kernel mapping of MSPA are also provided in this paper. The experimental results on 14 real-world datasets show our methods are superior to other state-of-the-art algorithms in discriminative DR tasks.
Le Yang 0007, Shiji Song, Shuang Li 0008, Yiming Chen 0005, C. L. Philip Chen
IEEE Trans. Cybern.2
2021 Spatially Adaptive Feature Refinement for Efficient Inference
abstract
Spatial redundancy commonly exists in the learned representations of convolutional neural networks (CNNs), leading to unnecessary computation on high-resolution features. In this paper, we propose a novel Spatially Adaptive feature Refinement (SAR) approach to reduce such superfluous computation. It performs efficient inference by adaptively fusing information from two branches: one conducts standard convolution on input features at a lower spatial resolution, and the other one selectively refines a set of regions at the original resolution. The two branches complement each other in feature learning, and both of them evoke much less computation than standard convolution. SAR is a flexible method that can be conveniently plugged into existing CNNs to establish models with reduced spatial redundancy. Experiments on CIFAR and ImageNet classification, COCO object detection and PASCAL VOC semantic segmentation tasks validate that the proposed SAR can consistently improve the network performance and efficiency. Notably, our results show that SAR only refines less than 40% of the regions in the feature representations of a ResNet for 97% of the samples in the validation set of ImageNet to achieve comparable accuracy with the original model, revealing the high computational redundancy in the spatial dimension of CNNs.
Yizeng Han, Gao Huang 0001, Shiji Song, Le Yang 0007, Haojun Jiang
IEEE Trans. Image Process.3
2021 Self-Attention-Based Temporary Curiosity in Reinforcement Learning Exploration
abstract
In many real-world scenarios, extrinsic rewards provided by the environment are sparse. An agent trained with classic reinforcement learning algorithm fails to explore these environments in a sufficient and effective way. To address this problem, the exploration bonus which derives from environmental novelty serves as intrinsic motivation for the agent. In recent years, curiosity-driven exploration is a mainstream approach to describe environmental novelty through prediction errors of dynamics models. Due to the expressive ability limitations of curiosity-based environmental novelty and the difficulty of finding appropriate feature space, most curiosity-driven exploration methods have the problem of overprotection against repetition. This problem can reduce the efficiency of exploration and lead the agent into a trap with local optimality. In this article, we propose a combination of persisting curiosity and temporary curiosity framework to deal with the problem of overprotection against repetition. We introduce the self-attention mechanism from the field of computer vision and propose a sequence-based self-attention mechanism for temporary curiosity generation. We compare our framework with some previous exploration methods in hard-exploration environments, provide a series of comprehensive analysis of the proposed framework and investigate the effect of the individual components of our method. The experimental results indicate that the proposed framework delivers superior performance than existing methods.
Hangkai Hu, Shiji Song, Gao Huang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Optimization-Based Control for Bearing-Only Target Search With a Mobile Vehicle
abstract
This article aims to design an optimization-based controller for a discrete-time Dubins vehicle to approach a target with unknown position as fast as possible by only using bearing measurements. To this end, we propose a bi-objective optimization problem, which jointly considers the performance of estimating the unknown target position and controlling the mobile vehicle to a known position, and then adopt a weighted sum method with normalization to solve it. The controller is given based on the solution of the optimization problem in ties with a least-square estimate of the target position. Moreover, the controller does not need the vehicle's global position information. Finally, the simulation results are included to validate the effectiveness of the proposed controller.
Zhuo Li 0011, Keyou You, Shiji Song, Anke Xue
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Input-to-State Stability of Nonlinear Systems Using Observer-Based Event-Triggered Impulsive Control
abstract
In this article, we are concerned with the input-to-state stability (ISS) problem for a class of nonlinear systems with exogenous disturbances, where the states of the system are not fully available. Based on the idea of event-triggered control (ETC) strategy, a novel event-triggered mechanism (ETM) is designed to reduce the burden of the communication and controller updating with guaranteed performance requirement. Correspondingly, an observer-based impulsive controller coupled with sample control is proposed such that the controlled system is ISS under the designed ETM. Moreover, the possible accumulations of triggered instants (i.e., Zeno behavior) in the proposed control strategy are excluded. The controller gains and ETM parameters are co-designed by solving linear matrix inequalities (LMIs). Finally, a numerical example is provided to illustrate the effectiveness of the obtained theoretical results.
Xiaodi Li 0001, Shiji Song
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Graph Embedding-Based Dimension Reduction With Extreme Learning Machine
abstract
Dimension reduction (DR)-based on extreme learning machine auto-encoder (ELM-AE) has achieved many successes in recent years. By minimizing the self-reconstruction error, the ELM-AE-based DR algorithms learn the compressed representations which facilitate the subsequent classification. However, the existing ELM-AEs only consider the DR problem in an unsupervised manner and ignore the valuable supervised information when these information is available. To find discriminative features of the original data, in this paper, we propose a graph embedding-based DR framework with ELM (GDR-ELM) for DR problems. Instead of self-reconstruction, the proposed GDR-ELM reconstructs all samples according to the weights in a graph matrix containing the supervised information. Furthermore, GDR-ELM can be stacked as building blocks to construct a multilayer framework like other ELM-AEs for more complicated representation learning tasks. Experiments on various datasets demonstrate the effectiveness of the proposed GDR-ELM and its multilayer framework.
Le Yang 0007, Shiji Song, Shuang Li 0008, Yiming Chen 0005, Gao Huang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Observer-Based Sliding Mode Control for Stabilization of Mismatched Disturbance Systems With or Without Time Delays
abstract
In this article, the stabilization problem of mismatched disturbance systems with or without time delays is studied via observed-based sliding mode control (SMC). Two kinds of SMC schemes for systems with time delay and without time delay are considered, respectively. When time delay is addressed in the system, an SMC strategy is structured via disturbance observer, where the unknown external disturbances are supposed to be generated by an exogenous dynamic. When time delay is not addressed, an SMC approach is designed in which the unknown external disturbances are assumed to tend to a constant steady state in infinite time. Sufficient conditions for stability of the corresponding sliding motion are derived by using Lyapunov–Krasovskii functional and Lyapunov function approach, respectively. Our results can be applied when the bound of the disturbances are unmeasured or unknown. Two simulation examples are shown to illustrate the proposed methods.
Yongshun Zhao, Xiaodi Li 0001, Shiji Song
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Resolution Adaptive Networks for Efficient Inference
abstract
Adaptive inference is an effective mechanism to achieve a dynamic tradeoff between accuracy and computational cost in deep networks. Existing works mainly exploit architecture redundancy in network depth or width. In this paper, we focus on spatial redundancy of input samples and propose a novel Resolution Adaptive Network (RANet), which is inspired by the intuition that low-resolution representations are sufficient for classifying “easy” inputs containing large objects with prototypical features, while only some “hard” samples need spatially detailed information. In RANet, the input images are first routed to a lightweight sub-network that efficiently extracts low-resolution representations, and those samples with high prediction confidence will exit early from the network without being further processed. Meanwhile, high-resolution paths in the network maintain the capability to recognize the “hard” samples. Therefore, RANet can effectively reduce the spatial redundancy involved in inferring high-resolution inputs. Empirically, we demonstrate the effectiveness of the proposed RANet on the CIFAR-10, CIFAR-100 and ImageNet datasets in both the anytime prediction setting and the budgeted batch classification setting.
Le Yang 0007, Yizeng Han, Shiji Song, Jifeng Dai, Gao Huang 0001
CVPR4
2020 Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification
abstract
The accuracy of deep convolutional neural networks (CNNs) generally improves when fueled with high resolution images. However, this often comes at a high computational cost and high memory footprint. Inspired by the fact that not all regions in an image are task-relevant, we propose a novel framework that performs efficient image classification by processing a sequence of relatively small inputs, which are strategically selected from the original image with reinforcement learning. Such a dynamic decision process naturally facilitates adaptive inference at test time, i.e., it can be terminated once the model is sufficiently confident about its prediction and thus avoids further redundant computation. Notably, our framework is general and flexible as it is compatible with most of the state-of-the-art light-weighted CNNs (such as MobileNets, EfficientNets and RegNets), which can be conveniently deployed as the backbone feature extractor. Experiments on ImageNet show that our method consistently improves the computational efficiency of a wide variety of deep models. For example, it further reduces the average latency of the highly efficient MobileNet-V3 on an iPhone XS Max by 20% without sacrificing accuracy. Code and pre-trained models are available at https://github.com/blackfeather-wang/GFNet-Pytorch.
Yulin Wang 0002, Kangchen Lv, Rui Huang 0012, Shiji Song, Le Yang 0007, Gao Huang 0001
NeurIPS4
2020 Collaborative learning with corrupted labels
Yulin Wang 0002, Rui Huang 0012, Gao Huang 0001, Shiji Song, Cheng Wu 0002
Neural Networks4
2020 Robust Scheduling of Hot Rolling Production by Local Search Enhanced Ant Colony Optimization Algorithm
abstract
Scheduling of a hot strip mill is an important decision problem in the steel manufacturing industry. Previous studies on the hot strip mill scheduling problem have mostly neglected the random factors in production. However, random variations in processing times are inevitable due to unpredictable delays and disturbances. In this article, we adopt a robust optimization approach to deal with the uncertainty in processing times. The advantage is that no assumption has to be made regarding the distribution of random data, and the obtained schedule will remain strictly feasible when the variations have not exceeded a predefined uncertainty set. First, a mixed-integer linear programming model is presented to formulate the robust scheduling problem. Then, a hybrid metaheuristic algorithm, which combines ant colony system (ACS) and enhanced local search, is proposed to provide an efficient solution to the problem. Finally, extensive computational experiments involving both randomly generated and real-world instances have been conducted to verify the effectiveness of the proposed algorithm. It is shown that the algorithm achieves optimality for small instances and outperforms two state-of-the-art metaheuristics when used to solve large instances.
Rui Zhang 0039, Shiji Song, Cheng Wu 0002
IEEE Trans. Ind. Informatics2
2020 A Graph Embedding Framework for Maximum Mean Discrepancy-Based Domain Adaptation Algorithms
abstract
Domain adaptation aims to deal with learning problems in which the labeled training data and unlabeled testing data are differently distributed. Maximum mean discrepancy (MMD), as a distribution distance measure, is minimized in various domain adaptation algorithms for eliminating domain divergence. We analyze empirical MMD from the point of view of graph embedding. It is discovered from the MMD intrinsic graph that, when the empirical MMD is minimized, the compactness within each domain and each class is simultaneously reduced. Therefore, points from different classes may mutually overlap, leading to unsatisfactory classification results. To deal with this issue, we present a graph embedding framework with intrinsic and penalty graphs for MMD-based domain adaptation algorithms. In the framework, we revise the intrinsic graph of MMD-based algorithms such that the within-class scatter is minimized, and thus, the new features are discriminative. Two strategies are proposed. Based on the strategies, we instantiate the framework by exploiting four models. Each model has a penalty graph characterizing certain similarity property that should be avoided. Comprehensive experiments on visual cross-domain benchmark datasets demonstrate that the proposed models can greatly enhance the classification performance compared with the state-of-the-art methods.
Yiming Chen 0005, Shiji Song, Shuang Li 0008, Cheng Wu 0002
IEEE Trans. Image Process.2
2020 Design of State-Dependent Switching Laws for Stability of Switched Stochastic Neural Networks With Time-Delays
abstract
We study the stability properties of switched stochastic neural networks (SSNNs) with time-varying delays whose subsystem is not necessarily stable. We introduce state-dependent switching (SDS) as a tool for stability analysis. Some SDS laws for asymptotic stability and p th moment exponentially stable are designed by employing Lyapunov-Krasovskii (L-K) functional and Lyapunov-Razumikhin (L-R) method, respectively. It is shown that the stability of SSNNs with time-varying delays composed of unstable subsystems can be achieved by using SDS law. The control gains in the designed SDS laws can be derived by solving the LMIs in derived stability criteria. Two numerical examples are provided to demonstrate the effectiveness of the proposed SDS laws.
Dan Yang 0013, Xiaodi Li 0001, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.3
2019 Soft Policy Gradient Method for Maximum Entropy Deep Reinforcement Learning
abstract
Maximum entropy deep reinforcement learning (RL) methods have been demonstrated on a range of challenging continuous tasks. However, existing methods either suffer from severe instability when training on large off-policy data or cannot scale to tasks with very high state and action dimensionality such as 3D humanoid locomotion. Besides, the optimality of desired Boltzmann policy set for non-optimal soft value function is not persuasive enough. In this paper, we first derive soft policy gradient based on entropy regularized expected reward objective for RL with continuous actions. Then, we present an off-policy actor-critic, model-free maximum entropy deep RL algorithm called deep soft policy gradient (DSPG) by combining soft policy gradient with soft Bellman equation. To ensure stable learning while eliminating the need of two separate critics for soft value functions, we leverage double sampling approach to making the soft Bellman equation tractable. The experimental results demonstrate that our method outperforms in performance over off-policy prior methods.
Wenjie Shi, Shiji Song
IJCAI2
2019 End-to-end sensorimotor control problems of AUVs with deep reinforcement learning
abstract
This paper studies on sensorimotor control problems of Autonomous Underwater Vehicles (AUVs) using deep reinforcement learning. We design an end-to-end learning architecture mapping original sensor input to continuous control output without referring to the dynamics of vehicles. To avoid difficult and noisy underwater localization, we implement the learning without knowing the positions of AUVs by proposing novel state encoder and reward shaping strategies. Two distinct underwater tasks, obstacle avoidance with sonar sensor and pipeline following with visual sensor, are simulated to validate the effectiveness of proposed architecture and strategies. For the latter, we test the learned policy on realistic images of underwater pipelines to check its generalization ability.
Hui Wu 0002, Shiji Song, Yachu Hsu, Keyou You, Cheng Wu 0002
IROS2
2019 Regularized Anderson Acceleration for Off-Policy Deep Reinforcement Learning
abstract
Model-free deep reinforcement learning (RL) algorithms have been widely used for a range of complex control tasks. However, slow convergence and sample inefficiency remain challenging problems in RL, especially when handling continuous and high-dimensional state spaces. To tackle this problem, we propose a general acceleration method for model-free, off-policy deep RL algorithms by drawing the idea underlying regularized Anderson acceleration (RAA), which is an effective approach to accelerating the solving of fixed point problems with perturbations. Specifically, we first explain how policy iteration can be applied directly with Anderson acceleration. Then we extend RAA to the case of deep RL by introducing a regularization term to control the impact of perturbation induced by function approximation errors. We further propose two strategies, i.e., progressive update and adaptive restart, to enhance the performance. The effectiveness of our method is evaluated on a variety of benchmark tasks, including Atari 2600 and MuJoCo. Experimental results show that our approach substantially improves both the learning speed and final performance of state-of-the-art deep RL algorithms.
Wenjie Shi, Shiji Song, Hui Wu 0002, Yachu Hsu, Cheng Wu 0002, Gao Huang 0001
NeurIPS2
2019 Implicit Semantic Data Augmentation for Deep Networks
abstract
In this paper, we propose a novel implicit semantic data augmentation (ISDA) approach to complement traditional augmentation techniques like flipping, translation or rotation. Our work is motivated by the intriguing property that deep networks are surprisingly good at linearizing features, such that certain directions in the deep feature space correspond to meaningful semantic transformations, e.g., adding sunglasses or changing backgrounds. As a consequence, translating training samples along many semantic directions in the feature space can effectively augment the dataset to improve generalization. To implement this idea effectively and efficiently, we first perform an online estimate of the covariance matrix of deep features for each class, which captures the intra-class semantic variations. Then random vectors are drawn from a zero-mean normal distribution with the estimated covariance to augment the training data in that class. Importantly, instead of augmenting the samples explicitly, we can directly minimize an upper bound of the expected cross-entropy (CE) loss on the augmented training set, leading to a highly efficient algorithm. In fact, we show that the proposed ISDA amounts to minimizing a novel robust CE loss, which adds negligible extra computational cost to a normal training procedure. Although being simple, ISDA consistently improves the generalization performance of popular deep models (ResNets and DenseNets) on a variety of datasets, e.g., CIFAR-10, CIFAR-100 and ImageNet. Code for reproducing our results are available at https://github.com/blackfeather-wang/ISDA-for-Deep-Networks.
Yulin Wang 0002, Xuran Pan, Shiji Song, Hong Zhang 0009, Gao Huang 0001, Cheng Wu 0002
NeurIPS3
2019 Distributed range-free localization via hierarchical nonconvex constrained optimization
Pei Xie, Keyou You, Shiji Song
Signal Process.3
2019 Domain Space Transfer Extreme Learning Machine for Domain Adaptation
abstract
Extreme learning machine (ELM) has been applied in a wide range of classification and regression problems due to its high accuracy and efficiency. However, ELM can only deal with cases where training and testing data are from identical distribution, while in real world situations, this assumption is often violated. As a result, ELM performs poorly in domain adaptation problems, in which the training data (source domain) and testing data (target domain) are differently distributed but somehow related. In this paper, an ELM-based space learning algorithm, domain space transfer ELM (DST-ELM), is developed to deal with unsupervised domain adaptation problems. To be specific, through DST-ELM, the source and target data are reconstructed in a domain invariant space with target data labels unavailable. Two goals are achieved simultaneously. One is that, the target data are input into an ELM-based feature space learning network, and the output is supposed to approximate the input such that the target domain structural knowledge and the intrinsic discriminative information can be preserved as much as possible. The other one is that, the source data are projected into the same space as the target data and the distribution distance between the two domains is minimized in the space. This unsupervised feature transformation network is followed by an adaptive ELM classifier which is trained from the transferred labeled source samples, and is used for target data label prediction. Moreover, the ELMs in the proposed method, including both the space learning ELM and the classifier, require just a small number of hidden nodes, thus maintaining low computation complexity. Extensive experiments on real-world image and text datasets are conducted and verify that our approach outperforms several existing domain adaptation methods in terms of accuracy while maintaining high efficiency.
Yiming Chen 0005, Shiji Song, Shuang Li 0008, Le Yang 0007, Cheng Wu 0002
IEEE Trans. Cybern.2
2019 Plume Tracing via Model-Free Reinforcement Learning Method
abstract
This paper studies the plume-tracing strategy for an autonomous underwater vehicle (AUV) in the deep-sea turbulent environment. The tracing problem is modeled as a partially observable Markov decision process with continuous state space and action space due to the spatio-temporal changes of environment. An long short-term memory-based reinforcement learning framework with full use of history information is proposed to generate a smooth strategy while the AUV interacting with the environment. Continuous temporal difference and deterministic policy gradient methods are employed to improve the strategy. To promote the performance of the algorithm, a supervised strategy generated by dynamic programming methods is utilized as transcendental knowledge of the agent. Historical searching trajectory's form and the exploration technology are specially designed to fit the algorithm. Simulation environments are established based on Reynolds-averaged Navier-Stokes equations and the effectiveness of the learned plume-tracing strategy is validated with simulation experiments.
Hangkai Hu, Shiji Song, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2019 Multi Pseudo Q-Learning-Based Deterministic Policy Gradient for Tracking Control of Autonomous Underwater Vehicles
abstract
This paper investigates trajectory tracking problem for a class of underactuated autonomous underwater vehicles (AUVs) with unknown dynamics and constrained inputs. Different from existing policy gradient methods which employ single actor critic but cannot realize satisfactory tracking control accuracy and stable learning, our proposed algorithm can achieve high-level tracking control accuracy of AUVs and stable learning by applying a hybrid actors-critics architecture, where multiple actors and critics are trained to learn a deterministic policy and action-value function, respectively. Specifically, for the critics, the expected absolute Bellman error-based updating rule is used to choose the worst critic to be updated in each time step. Subsequently, to calculate the loss function with more accurate target value for the chosen critic, Pseudo Q-learning, which uses subgreedy policy to replace the greedy policy in Q-learning, is developed for continuous action spaces, and Multi Pseudo Q-learning (MPQ) is proposed to reduce the overestimation of action-value function and to stabilize the learning. As for the actors, deterministic policy gradient is applied to update the weights, and the final learned policy is defined as the average of all actors to avoid large but bad updates. Moreover, the stability analysis of the learning is given qualitatively. The effectiveness and generality of the proposed MPQ-based deterministic policy gradient (MPQ-DPG) algorithm are verified by the application on AUV with two different reference trajectories. In addition, the results demonstrate high-level tracking control accuracy and stable learning of MPQ-DPG. Besides, the results also validate that increasing the number of the actors and critics will further improve the performance.
Wenjie Shi, Shiji Song, Cheng Wu 0002, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2019 Nonparametric Dimension Reduction via Maximizing Pairwise Separation Probability
abstract
In this brief, we propose a novel nonparametric supervised linear dimension reduction (SLDR) algorithm that extracts the features by maximizing the pairwise separation probability. The separation probability, as a new class separability measure, describes the generalization accuracy when we use the obtained features to train a linear classifier. Obtaining high-quality features, the proposed method avoids the overlaps between classes that are close to each other in the input space and improves the subsequent classification performance. Experiments on benchmark data sets show the superiority of the proposed algorithm over some other state-of-the-art SLDR methods.
Le Yang 0007, Shiji Song, Yanshang Gong, Gao Huang 0001, Cheng Wu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2019 Cross-Domain Extreme Learning Machines for Domain Adaptation
abstract
Extreme learning machines (ELMs), as “generalized” single hidden layer feedforward networks, have been proved to be effective and efficient for classification and regression problems. Traditional ELMs assume that the training and testing data are drawn from the same distribution, which however is often violated in real-world applications. In this paper, we propose a unified cross-domain ELM (CDELM) framework to address domain adaptation problems, in which the distributions of training data (source domain) and testing data (target domain) are different but related. CDELM not only fully leverages labeled source data and unlabeled target data simultaneously to construct an adaptive target classifier but also maintains the computational efficiency of ELMs. Specifically, CDELM adapts the source classifier to target domain by matching the projected means of both domains, and explores the structure property of target domain by using manifold regularization to make the final classifier more adaptable to target data. Based on the framework, two algorithms CDELM-M and CDELM-C are proposed, which aim at minimizing the marginal and conditional distribution distance between source and target domains, respectively. Moreover, CDELM-C can further enhance the classification accuracy by multiple iterations. Comprehensive experimental studies on artificial datasets and public text and image datasets demonstrate that both CDELM-M and CDELM-C are competitive with several state-of-the-art domain adaptation learning methods in terms of the classification accuracy and efficiency.
Shuang Li 0008, Shiji Song, Gao Huang 0001, Cheng Wu 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Depth Control of Model-Free AUVs via Reinforcement Learning
abstract
In this paper, we consider depth control problems of an autonomous underwater vehicle (AUV) for tracking the desired depth trajectories. Due to the unknown dynamical model and the coupling between surge and yaw motions of the AUV, the problems cannot be effectively solved by most of the model-based or proportional-integral-derivative like controllers. To this purpose, we formulate the depth control problems of the AUV as continuous-state, continuous-action Markov decision processes under unknown transition probabilities. Based on the deterministic policy gradient theorem and neural network approximation, we propose a model-free reinforcement learning (RL) algorithm that learns a state-feedback controller from sampled trajectories of the AUV. To improve the performance of the RL algorithm, we further propose a batch-learning scheme through replaying previous prioritized trajectories. We illustrate with simulations that our model-free method is even comparable to the model-based controllers. Moreover, we validate the effectiveness of the proposed RL algorithm on a seafloor data set sampled from the South China Sea.
Hui Wu 0002, Shiji Song, Keyou You, Cheng Wu 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2018 High-Level Tracking of Autonomous Underwater Vehicles Based on Pseudo Averaged Q-Learning
abstract
In this paper, we investigate the trajectory tracking problem of underactuated autonomous underwater vehicles (AUVs) with input saturation. Our proposed model-free algorithm can realize high-level tracking control and stable learning by employing a novel actors-critics architecture, where a critic and multiple actors are learned to estimate the action-value function and deterministic policy, respectively. For the critic, Pseudo Averaged Q-learning, which is a simple extension to Q-learning, is proposed to calculate the target value, specifically, the action-value of next state is obtained by maximizing the average over the last multiple previous learned action-value estimates among all actors. As for the actors, deterministic policy gradient is applied to update the weights. The effectiveness and performance of the proposed Pseudo Averaged Q-learning based deterministic policy gradient (PAQ-DPG) algorithm is verified by implementation to an underactuated AUV. And the results demonstrate high-level tracking control accuracy and stability of learning of PAQ-DPG algorithm. Besides, under our proposed actors-critics framework, increasing the number of actors will further improve the performance.
Wenjie Shi, Shiji Song
SMC2
2018 Transductive Transfer Learning Based on Broad Learning System
abstract
The latest proposed Broad Learning System (BLS) demonstrates an efficient and effective learning capability in many machine learning problems. In this paper, we apply the BLS to address transductive transfer learning problems, where the training (source) and test (target) data are drawn from the different but related distributions, which is a.k.a domain adaptation. We aim at learning from source data a well performing classifier on a different (but related) target data. A unified domain adaptation framework based on the BLS is developed for improving its transfer learning capability without loss of the computational efficiency. Two algorithms including BLS based source domain adaptation (BLS-SDA) and BLS based target domain adaptation (BLS-TDA) are proposed under this framework. Experiments on benchmark datasets show that our approach outperforms several existing domain adaptation methods while maintains high efficiency.
Le Yang 0007, Shiji Song, C. L. Philip Chen
SMC2
2018 Impulsive control of unstable neural networks with unbounded time-varying delays
Xiaodi Li 0001, Shiji Song, Jianhong Wu
Sci. China Inf. Sci.2
2018 Laplacian twin extreme learning machine for semi-supervised classification
Shuang Li 0008, Shiji Song, Yihe Wan
Neurocomputing2
2018 Exact Algorithms for Distributionally β-Robust Machine Scheduling with Uncertain Processing Times
abstract
The β-robust machine scheduling has attracted increasing attention as an effective method to hedge against uncertainty. However, existing β-robust scheduling models rely on the normality assumption of uncertain parameters, and existing solution methods are based on branch and bound, which cannot solve problems of 45 jobs within 3,600 seconds. This paper proposes distributionally β-robust scheduling (DRS) models to handle uncertain processing times. The DRS models only require the lower bound, mean, and covariance information of processing times, and have the capability of handling both single and parallel machine problems. Another key contribution of this paper is to devise efficient parametric search (PS) methods for the DRS models. Specifically, we show that there exists a parameterized assignment problem (PAP), such that its optimal solutions are also optimal for the original problem. The proposed methods only need to perform a one-dimensional PS and solve a series of PAPs. We further propose a bidirectional PS to reduce the number of PAPs needed to be solved, and we design a speedup shortest augmentation path algorithm for these PAPs. Experimental results on both single and identical parallel machine problems show that the improved PS method outperforms existing algorithms by more than three orders of magnitude improvement in computation time for problems of 45 jobs, and it is able to solve problems of 500 jobs within 0.5 seconds.
Zuo-Jun Max Shen, Shiji Song
INFORMS J. Comput.3
2018 Layer-wise domain correction for unsupervised domain adaptation
abstract
Deep neural networks have been successfully applied to numerous machine learning tasks because of their impressive feature abstraction capabilities. However, conventional deep networks assume that the training and test data are sampled from the same distribution, and this assumption is often violated in real-world scenarios. To address the domain shift or data bias problems, we introduce layer-wise domain correction (LDC), a new unsupervised domain adaptation algorithm which adapts an existing deep network through additive correction layers spaced throughout the network. Through the additive layers, the representations of source and target domains can be perfectly aligned. The corrections that are trained via maximum mean discrepancy, adapt to the target domain while increasing the representational capacity of the network. LDC requires no target labels, achieves state-of-the-art performance across several adaptation benchmarks, and requires significantly less training time than existing adaptation methods.
Shuang Li 0008, Shiji Song, Cheng Wu 0002
Frontiers Inf. Technol. Electron. Eng.2
2018 Real-time localization in wireless sensor network with multimedia applications
Shiji Song
Multim. Tools Appl.4
2018 Domain Invariant and Class Discriminative Feature Learning for Visual Domain Adaptation
abstract
Domain adaptation manages to build an effective target classifier or regression model for unlabeled target data by utilizing the well-labeled source data but lying different distributions. Intuitively, to address domain shift problem, it is crucial to learn domain invariant features across domains, and most existing approaches have concentrated on it. However, they often do not directly constrain the learned features to be class discriminative for both source and target data, which is of vital importance for the final classification. Therefore, in this paper, we put forward a novel feature learning method for domain adaptation to construct both domain invariant and class discriminative representations, referred to as DICD. Specifically, DICD is to learn a latent feature space with important data properties preserved, which reduces the domain difference by jointly matching the marginal and class-conditional distributions of both domains, and simultaneously maximizes the inter-class dispersion and minimizes the intra-class scatter as much as possible. Experiments in this paper have demonstrated that the class discriminative properties will dramatically alleviate the cross-domain distribution inconsistency, which further boosts the classification performance. Moreover, we show that exploring both domain invariance and class discriminativeness of the learned representations can be integrated into one optimization framework, and the optimal solution can be derived effectively by solving a generalized eigen-decomposition problem. Comprehensive experiments on several visual cross-domain classification tasks verify that DICD can outperform the competitors significantly.
Shuang Li 0008, Shiji Song, Gao Huang 0001, Zhengming Ding, Cheng Wu 0002
IEEE Trans. Image Process.2
2018 Robust Shortest Path Problem With Distributional Uncertainty
abstract
Routing service considering uncertainty is at the core of intelligent transportation systems and has attracted increasing attention. Existing stochastic shortest path models require the exact probability distributions of travel times and usually assume that they are independent. However, the distributions are often unavailable or inaccurate due to insufficient data, and correlation of travel times over different links has been observed. This paper presents a robust shortest path (RSP) model that only requires partial distribution information of travel times, including the support set, mean, variance, and correlation matrix. We introduce a concept of robust mean-excess travel time to hedge against the risk from both the uncertainty of the random travel times and the uncertainty in their distributions. To solve the RSP problem, an equivalent dual formulation is derived and used to design tight lower and upper bound approximation methods, which adopt the scenario approach and semi-definite programming approach, respectively. To solve large problems, we further propose an efficient primal approximation method, which only needs to solve two deterministic shortest path problems and a mean-standard deviation shortest path problem, and analyze its approximation performance. Experiments validate the tightness of the proposed bounds and demonstrate the impact of uncertainty on the relative benefit and cost of robust paths.
Shiji Song, Zuo-Jun Max Shen, Cheng Wu 0002
IEEE Trans. Intell. Transp. Syst.2
2017 Twin extreme learning machines for pattern classification
Yihe Wan, Shiji Song, Gao Huang 0001, Shuang Li 0008
Neurocomputing2
2017 A multi-objective artificial bee colony algorithm for parallel batch-processing machine scheduling in fabric dyeing processes
Rui Zhang 0039, Pei-Chann Chang, Shiji Song, Cheng Wu 0002
Knowl. Based Syst.3
2017 Prediction Reweighting for Domain Adaptation
abstract
There are plenty of classification methods that perform well when training and testing data are drawn from the same distribution. However, in real applications, this condition may be violated, which causes degradation of classification accuracy. Domain adaptation is an effective approach to address this problem. In this paper, we propose a general domain adaptation framework from the perspective of prediction reweighting, from which a novel approach is derived. Different from the major domain adaptation methods, our idea is to reweight predictions of the training classifier on testing data according to their signed distance to the domain separator, which is a classifier that distinguishes training data (from source domain) and testing data (from target domain). We then propagate the labels of target instances with larger weights to ones with smaller weights by introducing a manifold regularization method. It can be proved that our reweighting scheme effectively brings the source and target domains closer to each other in an appropriate sense, such that classification in target domain becomes easier. The proposed method can be implemented efficiently by a simple two-stage algorithm, and the target classifier has a closed-form solution. The effectiveness of our approach is verified by the experiments on artificial datasets and two standard benchmarks, a visual object recognition task and a cross-domain sentiment analysis of text. Experimental results demonstrate that our method is competitive with the state-of-the-art domain adaptation algorithms.
Shuang Li 0008, Shiji Song, Gao Huang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2017 Dimension Reduction by Minimum Error Minimax Probability Machine
abstract
Dimension reduction is frequently adopted as a data preprocessing technique to facilitate data visualization, interpretation, and classification. Traditional dimension reduction methods such as linear discriminant analysis focus on maximizing the overall discrimination between all classes, which may be easily affected by outliers. To overcome this disadvantage, this paper proposes a novel method for multiclass dimension reduction, named dimension reduction by minimum error minimax probability machine (DR-MEMPM). It elaborately ensures that each pair of classes is well separated in the projected subspace by utilizing the separation probability between different pairwise classes. Therefore, it can put more emphasis on those less distinguishable classes, and the learned projection will not be dominated by some “outlier” classes which lie far away from other classes. We evaluate the proposed DR-MEMPM on a number of synthetic and real-world data sets, and show that it outperforms other state-of-the-art dimension reduction methods in terms of visual intuition and classification accuracy, especially when the distances between classes are unevenly distributed.
Shiji Song, Yanshang Gong, Gao Huang 0001, Guang-Bin Huang
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Efficient and Rapid Machine Learning Algorithms for Big Data and Dynamic Varying Systems
abstract
With the exponential growth of data and complexity of systems, fast machine learning/artificial intelligence and computational intelligence techniques are highly required. Many conventional computational intelligence techniques face bottlenecks in learning (e.g., intensive human intervention and convergence time) [item 1) in the Appendix]. However, efficient learning algorithms alternatively offer significant benefits including fast learning speed, ease of implementation, and minimal human intervention. The need for efficient and fast implementation of machine learning techniques in big data and dynamic varying systems poses many research challenges. This special issue highlights some latest development in the related areas.
Fuchun Sun 0001, Guang-Bin Huang, Q. M. Jonathan Wu, Shiji Song, Donald C. Wunsch II
IEEE Trans. Syst. Man Cybern. Syst.4
2016 A source-seeking strategy for an autonomous underwater vehicle via on-line field estimation
abstract
This paper studies the problem of using an autonomous underwater vehicle (AUV) to seek the source of some signal in the underwater environment. To avoid the redundant travels for getting the local gradient in existing methods, we propose a novel source-seeking strategy in which the AUV keeps updating an estimated field model using measurements along its gradient-climbing path. Based on the convergence results of the sequential least-squares field estimation algorithm and the path planning method by nonlinear programming, the AUV will finally reach a maximum of the field which indicates a potential source. The way-point tracking control of the AUV is realized by a nonlinear model predictive control (NMPC) scheme embedded with the line-of-sight (LOS) guidance law. The effectiveness and efficiency of the proposed source-seeking strategy is validated in the simulation experiment with a real AUV model.
Xiaodong Ai, Keyou You, Shiji Song
ICARCV3
2016 Unsupervised learning of Dirichlet process mixture models with missing data
Xunan Zhang, Shiji Song, Keyou You
Sci. China Inf. Sci.2
2016 Parallel Machine Scheduling Under Time-of-Use Electricity Prices: New Models and Optimization Approaches
abstract
The industrial sector is one of the largest energy consumers in the world. To alleviate the grid's burden during peak hours, time-of-use (TOU) electricity pricing has been implemented in many countries around the globe to encourage manufacturers to shift their electricity usage from peak periods to off-peak periods. In this paper, we study the unrelated parallel machine scheduling problem under a TOU pricing scheme. The objective is to minimize the total electricity cost by appropriately scheduling the jobs such that the overall completion time does not exceed a predetermined production deadline. To solve this problem, two solution approaches are presented. The first approach models the problem with a new time-interval-based mixed integer linear programming formulation. In the second approach, we reformulate the problem using Dantzig-Wolfe decomposition and propose a column generation heuristic to solve it. Computational experiments are conducted under different TOU settings and the results confirm the effectiveness of the proposed methods. Based on the numerical results, we provide some practical suggestions for decision makers to help them in achieving a good balance between the productivity objective and the energy cost objective.
Jianya Ding, Shiji Song, Rui Zhang 0039, Raymond Chiong, Cheng Wu 0002
IEEE Trans Autom. Sci. Eng.2
2015 Maximin Separation Probability Clustering
abstract
This paper proposes a new approach for discriminative clustering. The intuition is, for a good clustering, one should be able to learn a classifier from the clustering labels with high generalization accuracy. Thus we define a novel metric to evaluate the quality of a clustering labeling, named Minimum Separation Probability (MSP), which is a lower bound of the generalization accuracy of a classifier learnt from the clustering labeling. We take MSP as the objective to maximize and propose our approach Maximin Separation Probability Clustering (MSPC), which has several attractive properties, such as invariance to anisotropic feature scaling and intuitive probabilistic explanation for clustering quality. We present three efficient optimization strategies for MSPC, and analyze their interesting connections to existing clustering approaches, such as maximum margin clustering (MMC) and discriminative k-means. Empirical results on real world data sets verify that MSP is a robust and effective clustering quality measure. It is also shown that the proposed algorithms compare favorably to state-of-the-art clustering algorithms in both accuracy and efficiency.
Gao Huang 0001, Shiji Song, Zheng Chen 0001
AAAI3
2015 A Reduction of the Elastic Net to Support Vector Machines with an Application to GPU Computing
abstract
Algorithmic reductions are one of the corner stones of theoretical computer science. Surprisingly, to-date, they have only played a limited role in machine learning. In this paper we introduce a formal and practical reduction between two of the most widely used machine learning algorithms: from the Elastic Net (and the Lasso as a special case) to the Support Vector Machine. First, we derive the reduction and summarize it in only 11 lines of MATLAB. Then, we demonstrate its high impact potential by translating recent advances in parallelizing SVM solvers directly to the Elastic Net. The resulting algorithm is a parallel solver for the Elastic Net (and Lasso) that naturally utilizes GPU and multi-core CPUs. We evaluate it on twelve real world data sets, and show that it yields identical results as the popular (and highly optimized) glmnet implementation but is up-to two orders of magnitude faster.
Shiji Song, Jacob R. Gardner, Kilian Q. Weinberger, Yixin Chen 0001
AAAI3
2015 A novel Block-shifting simulated annealing algorithm for the no-wait flowshop scheduling problem
abstract
This paper proposes a Block-shifting Simulated Annealing (BSA) algorithm for the no-wait flowshop scheduling problem (NWFSP) to minimize makespan. The proposed algorithm makes use of the objective incremental properties of NWFSP and embeds a block-shifting operator based on k-insertion moves into the algorithm framework of simulated annealing. A major advantage of the BSA algorithm lies in its easy implementation since it does not involve sophisticated evolutionary strategy and parameter tuning process. In addition to its simplicity, BSA is shown to be very effective. Through experimental comparisons, the effectiveness of the block-shifting operator is clearly revealed. In addition, the BSA algorithm is proved to be more effective and robust than the state-of-the-art algorithms for solving the NWFSP.
Jianya Ding, Shiji Song, Rui Zhang 0039, Cheng Wu 0002
CEC2
2015 Solving a Multi-Level Capacitated Lot Sizing Problem with Random Demand via a Fix-and-Optimize heuristic
abstract
In supply chain management, the multi-level multi-product capacitated lot sizing problem (MLCLSP) plays a critical role in operational production systems, where limited resources and time phases are considered. In this paper, we present a stochastic version of MLCLSP subject to capacity constraints. In order to minimize the total expected cost of MLCLSP, a production schedule has to be determined for random demand and a new backlog-oriented δ-service-level measure should be met. We apply scenario method to approximate the stochastic MLCLSP and lead to a new MLCLSP-SCN model. For the purpose of determining production schedule efficiently, robustly and stably, a new heuristic called Fix-and-Optimize heuristic is presented handling random demand. A numerical analysis based on a set of artificial test instances is used to evaluate the relative performance of the heuristic. We further present the numerical results of dynamic programming approach and heuristic genetic algorithm as benchmarks. All the numerical study suggests that our heuristic provides high-quality solutions and the computational effort is moderate.
Liuxi Li, Shiji Song
CEC2
2015 Dimension Reduction by Maximizing Pairwise Discriminations
abstract
Dimension reduction is an important pre-processing technique for high-dimensional data analysis. In this paper, we consider the process of linear dimension reduction (LDR) in multiclass problems. We propose a novel feature extraction method based on Minimax Probability Machine (MPM), named MPMbased Dimension Reduction (DR-MPM). Its objective naturally combines the discriminative information of all the class pairs and each pair of classes is well separated in the projected subspace. The algorithm is robust in the sense that it is insensitive to 'outlier' classes which lie far away from other classes. We evaluate DR-MPM on a number of synthetic and real-world data sets, and show that it outperforms other state-of-art feature extraction methods in terms of visual intuition and classification accuracy, especially when the distances between classes are unevenly distributed.
Yanshang Gong, Shiji Song, Gao Huang 0001
SMC2
2015 Incremental Extreme Learning Machine Based on Cascade Neural Networks
abstract
This paper extends extreme learning machine (ELM) for multi-layer cascade neural networks. We reformulate the cascade neural networks as a linear-in-the-parameters model, and propose a novel constructive training algorithm motivated by the efficient incremental ELM. The orthogonal least squares (OLS) is introduced to derive a new criterion for evaluating candidate hidden units, which avoids the computation of Moore-Penrose generalized inverse in the training process. Moreover, the calculation of output weights can be greatly simplified. Besides its efficiency, we show that the proposed evaluation function can effectively identify optimal candidate unit which leads to maximum error (sum of squared errors, SSE) reduction of the network. As a result, the proposed algorithm tends to yield smaller network with better generalization performance compared to traditional ELM. The effectiveness of the proposed algorithm on classification and regression problems is demonstrated by experimental results on several real-world datasets.
Yihe Wan, Shiji Song, Gao Huang 0001
SMC2
2015 Dual active set method for support vector machines under multi-constraint activation
Shiji Song, Keyou You
Neurocomputing2
2015 Unsupervised neighborhood component analysis for clustering
Chen Qin, Shiji Song, Gao Huang 0001
Neurocomputing2
2015 Variable exponential neighborhood search for the long chain design problem
Shiji Song, Cheng Wu 0002, Wenjun Yin
Neurocomputing2
2015 Efficient Lasso training from a geometrical perspective
Shiji Song, Gao Huang 0001, Cheng Wu 0002
Neurocomputing2
2015 Trends in extreme learning machines: A review
Gao Huang 0001, Guang-Bin Huang, Shiji Song, Keyou You
Neural Networks3
2015 Discriminative clustering via extreme learning machine
Gao Huang 0001, Tianchi Liu 0001, Zhiping Lin 0001, Shiji Song, Cheng Wu 0002
Neural Networks5
2014 Minimizing makespan for a no-wait flowshop using tabu mechanism improved iterated greedy algorithm
abstract
This paper proposes a tabu mechanism improved iterated greedy (TMIIG) algorithm to solve the no-wait flow-shop scheduling problem with makespan criterion. The motivation of seeking for further improvement in the iterated greedy (IG) algorithm framework is based on the observation that the construction phase of the original IG algorithm may lead to repeated search when applying the insertion neighborhood search. To overcome the drawback, we modified the IG algorithm by a tabu-based reconstruction strategy to enhance its exploitation ability. A powerful neighborhood search method which involves insert, swap, and double-insert moves is then applied to obtain better soluions from the reconstructed solution in the previous step. Numerical computations verified the advantages of utilizing the new reconstruction scheme. In addition, comparisons with other high-performing algorithms demonstrated the effectiveness and robustness of the proposed algorithm.
Jianya Ding, Shiji Song, Rui Zhang 0039, Cheng Wu 0002
IEEE Congress on Evolutionary Computation2
2014 Transductive Minimax Probability Machine
Gao Huang 0001, Shiji Song, Zhixiang Eddie Xu, Kilian Q. Weinberger
ECML/PKDD (1)2
2014 Non-linear neighborhood component analysis based on constructive neural networks
abstract
In this paper, we propose a novel non-linear supervised metric learning algorithm. The algorithm combines the neighborhood component analysis method with constructive neural networks which gradually increase the network size during the training process. The network aims to maximize a stochastic variant of the leave-one-out K-nearest neighbor (KNN) score on the training set. In this way, the proposed algorithm learns a nonlinear metric for KNN classification, overcoming the limitations of traditional metric learning algorithms which are only capable of learning linear transformations. Therefore, the proposed method is more flexible and powerful in transforming data than its linear counterpart. Moreover, it can also learn a low-dimensional non-linear mapping for visualization and fast classification. We validate our method on several benchmark datasets both for metric learning and dimensionality reduction, and the results demonstrate the competitiveness of the proposed approach.
Chen Qin, Shiji Song, Gao Huang 0001
SMC2
2014 Semi-Supervised and Unsupervised Extreme Learning Machines
abstract
Extreme learning machines (ELMs) have proven to be efficient and effective learning mechanisms for pattern classification and regression. However, ELMs are primarily applied to supervised learning problems. Only a few existing research papers have used ELMs to explore unlabeled data. In this paper, we extend ELMs for both semi-supervised and unsupervised tasks based on the manifold regularization, thus greatly expanding the applicability of ELMs. The key advantages of the proposed algorithms are as follows: 1) both the semi-supervised ELM (SS-ELM) and the unsupervised ELM (US-ELM) exhibit learning capability and computational efficiency of ELMs; 2) both algorithms naturally handle multiclass classification or multicluster clustering; and 3) both algorithms are inductive and can handle unseen data at test time directly. Moreover, it is shown in this paper that all the supervised, semi-supervised, and unsupervised ELMs can actually be put into a unified framework. This provides new perspectives for understanding the mechanism of random feature mapping, which is the key concept in ELM theory. Empirical study on a wide range of data sets demonstrates that the proposed algorithms are competitive with the state-of-the-art semi-supervised or unsupervised learning algorithms in terms of accuracy and efficiency.
Gao Huang 0001, Shiji Song, Jatinder N. D. Gupta, Cheng Wu 0002
IEEE Trans. Cybern.2
2013 A fast iterative single data approach to training unconstrained least squares support vector machines
Shiji Song
Neurocomputing2
2013 Kernelized LARS-LASSO for constructing radial basis function neural networks
Shiji Song, Cheng Wu 0002, Gao Huang 0001
Neural Comput. Appl.2
2013 A second order cone programming approach for semi-supervised learning
Gao Huang 0001, Shiji Song, Jatinder N. D. Gupta, Cheng Wu 0002
Pattern Recognit.2
2013 Impulsive Control for Existence, Uniqueness, and Global Stability of Periodic Solutions of Recurrent Neural Networks With Discrete and Continuously Distributed Delays
abstract
In this paper, a class of recurrent neural networks with discrete and continuously distributed delays is considered. Sufficient conditions for the existence, uniqueness, and global exponential stability of a periodic solution are obtained by using contraction mapping theorem and stability theory on impulsive functional differential equations. The proposed method, which differs from the existing results in the literature, shows that network models may admit a periodic solution which is globally exponentially stable via proper impulsive control strategies even if it is originally unstable or divergent. Two numerical examples and their computer simulations are offered to show the effectiveness of our new results.
Xiaodi Li 0001, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.2
2012 A two-stage hybrid particle swarm optimization algorithm for the stochastic job shop scheduling problem
Rui Zhang 0039, Shiji Song, Cheng Wu 0002
Knowl. Based Syst.2
2012 A hybrid genetic algorithm for two-stage multi-item inventory system with stochastic demand
Shiji Song, Heming Zhang 0001, Cheng Wu 0002, Wenjun Yin
Neural Comput. Appl.2
2012 Improved conjugate gradient implementation for least squares support vector machines
Shiji Song
Pattern Recognit. Lett.2
2012 Robust Support Vector Regression for Uncertain Input and Output Data
abstract
In this paper, a robust support vector regression (RSVR) method with uncertain input and output data is studied. First, the data uncertainties are investigated under a stochastic framework and two linear robust formulations are derived. Linear formulations robust to ellipsoidal uncertainties are also considered from a geometric perspective. Second, kernelized RSVR formulations are established for nonlinear regression problems. Both linear and nonlinear formulations are converted to second-order cone programming problems, which can be solved efficiently by the interior point method. Simulation demonstrates that the proposed method outperforms existing RSVRs in the presence of both input and output data uncertainties.
Gao Huang 0001, Shiji Song, Cheng Wu 0002, Keyou You
IEEE Trans. Neural Networks Learn. Syst.2
2011 Two Techniques to Improve the NEH Algorithm for Flow-Shop Scheduling Problems
Gengcheng Liu, Shiji Song
ICIC (2)2
2011 Truncation error calculation based on Richardson extrapolation for variable-step collaborative simulation
Heming Zhang 0001, Silv Liang, Shiji Song, Hongwei Wang 0001
Sci. China Inf. Sci.3
2010 Generalized gradient projection neural networks for nonsmooth optimization problems
Shiji Song
Sci. China Inf. Sci.2
2008 A Simulated Annealing Based Beam Search Algorithm for the Flow-Shop Scheduling Problem
abstract
Beam search algorithm, as an adaptation of branch and bound method, is regarded as one of the effective approaches in solving combinational optimization problems. In this paper, a new beam search algorithm for the large-scale permutation flow shop scheduling problem (FSP) is proposed. A new branching scheme is addressed and compared with the traditional branching scheme. With the new branching scheme, the number of partial schedules in the search tree can be greatly reduced. Based on a simple simulated annealing algorithm, partial schedules are globally evaluated. Numerical experiments show that good solutions of large-scale FSPs could be found with the proposed algorithm in a short time.
Shiji Song
Int. J. Pattern Recognit. Artif. Intell.2
2008 Acceleration-based Dopplerlet transform - Part II: Implementations and applications to passive motion parameter estimation of moving sound source
Hongxing Zou, Lin Qiao, Shiji Song, Yanda Li
Signal Process.4
2008 Acceleration-based Dopplerlet transform - Part I: Theory
Hongxing Zou, Shiji Song, Yanda Li
Signal Process.2
2006 Detecting Link Spam Using Temporal Information
abstract
How to effectively protect against spam on search ranking results is an important issue for contemporary web search engines. This paper addresses the problem of combating one major type of web spam: 'link spam.' Most of the previous work on anti link spam managed to make use of one snapshot of web data to detect spam, and thus it did not take advantage of the fact that link spam tends to result in drastic changes of links in a short time period. To overcome the shortcoming, this paper proposes using temporal information on links in detection of link spam, as well as other information. Specifically, it defines temporal features such as in-link growth rate (IGR) and in-link death rate (IDR) in a spam classification model (i.e., SVM). Experimental results on web domain graph data show that link spam can be successfully detected with the proposed method.
Guoyang Shen, Bin Gao 0001, Tie-Yan Liu, Shiji Song, Hang Li 0001
ICDM5
2006 A Neural Network Model for Non-smooth Optimization over a Compact Convex Subset
Shiji Song, Zifang Du
ISNN (1)2
2006 Differential Inclusions-Based Neural Networks for Nonsmooth Convex Optimization on a Closed Convex Subset
Shiji Song, Xiaohong Guan
ISNN (1)1
2006 Subgradient-based feedback neural networks for non-differentiable convex optimization problems
Shiji Song
Sci. China Ser. F Inf. Sci.2
2003 On stability of delayed cellular neural networks with sigmoid output functions
Yaru Mo, Xiaoping Xue 0001, Shiji Song
Sci. China Ser. F Inf. Sci.3
2002 Reverse triple I method of fuzzy reasoning
Shiji Song
Sci. China Ser. F Inf. Sci.1
2000 Global existence of solutions to fuzzy differential equations
Shiji Song, Chunbo Feng
Fuzzy Sets Syst.1
2000 Existence and uniqueness of solutions to Cauchy problem of fuzzy differential equations
Shiji Song, Congxin Wu
Fuzzy Sets Syst.1
1999 Existence and comparison theorems to Volterra fuzzy integral equation in (En, D)1
Shiji Song, Qin Yu Liu, Qi-chun Xu
Fuzzy Sets Syst.1
1998 On the basic solutions to the generalized fuzzy integral equation
Congxin Wu, Shiji Song
Fuzzy Sets Syst.2
1997 The convergence of (DG) fuzzy integrals
Shiji Song, Peilin Shi
Fuzzy Sets Syst.1