EDBT 2026 Demo / reviewers in the wild / expert
Shiyu Jin
dblp:202/7052
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain-Adaptive Mamba for Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image classification aims to identify a new scene in target domain via learned knowledge from source domain using limited training samples. Existing cross-scene alignment approaches focus on aligning the global feature distribution between the source and target domains while overlooking the fine-grained alignment at different levels. Moreover, they mainly use Transformer architectures to model long-range dependencies across different channels but confront efficiency challenges due to their quadratic complexity, which limits classification performance in unsupervised domain adaptation tasks. To address these issues, a new domain-adaptive Mamba (DAMamba) is proposed for cross-scene hyperspectral image classification. First, a spectral-spatial Mamba is developed to extract high-order semantic features from the input data. Then, a domain-invariant prototype alignment method is proposed from three perspectives, i.e., intra-domain, inter-domain, and mini-batch, to produce reliable pseudo-labels and mitigate the spectral shift between the source and target domains. Finally, a fully connected layer is applied to the aligned features in the target domain to obtain the final classification results. Extensive evaluations across diverse cross-scene datasets demonstrate that our DAMamba outperforms existing state-of-the-art methods in classification accuracy and computing time. The code of this paper is available at https://github.com/PuhongDuan/DAMamba. Puhong Duan, Shiyu Jin, Xiaotian Lu, Lianhui Liang, Xudong Kang, Antonio Plaza |
IEEE Trans. Image Process. | 2 |
| 2025 | Exploiting Text Semantics for Few and Zero Shot Node Classification on Text-attributed GraphabstractText-attributed graph (TAG) provides a text description for each graph node, and few- and zero-shot node classification on TAGs have many applications in fields such as academia and social networks. Existing work utilizes various graph-based augmentation techniques to train the node and text embeddings, while text-based augmentations are largely unexplored. In this paper, we propose Text Semantics Augmentation (TSA) to improve accuracy by introducing more text semantic supervision signals. Specifically, we design two augmentation techniques, i.e., positive semantics matching and negative semantics contrast, to provide more reference texts for each graph node or text description. Positive semantic matching retrieves texts with similar embeddings to match with a graph node. Negative semantic contrast adds a negative prompt to construct a text description with the opposite semantics, which is contrasted with the original node and text. We evaluate TSA on 5 datasets and compare with 13 state-of-the-art baselines. The results show that TSA consistently outperforms all baselines, and its accuracy improvements over the best-performing baseline are usually over 5%. The code is at https://github.com/wyx11112/TSA. Yuxiang Wang 0013, Xiao Yan 0002, Shiyu Jin, Quanqing Xu, Chuang Hu, Yuanyuan Zhu 0001, Bo Du 0001, Jia Wu 0001, Jiawei Jiang 0001 |
IJCAI | 3 |
| 2024 | VIHE: Virtual In-Hand Eye Transformer for 3D Robotic ManipulationabstractIn this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong inductive bias for effectively recognizing the correct pose for the hand, especially for challenging high-precision tasks such as peg insertion. On 18 manipulation tasks in RLBench simulated environments, VIHE achieves a new state-of-the-art, with a 12% absolute improvement, increasing from 65% to 77% over the existing state-of-the-art model using 100 demonstrations per task. In real-world scenarios, VIHE can learn manipulation tasks with just a handful of demonstrations, highlighting its practical utility. Videos and code implementation can be found at our project site: https://vihe-3d.github.io. Weiyao Wang 0002, Shiyu Jin, Gregory D. Hager, Liangjun Zhang |
IROS | 3 |
| 2024 | RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language ModelabstractRecent advances in Large Language Models (LLMs) have showcased their remarkable reasoning capabilities, making them influential across various fields. However, in robotics, their use has primarily been limited to manipulation planning tasks due to their inherent textual output. This paper addresses this limitation by investigating the potential of adopting the reasoning ability of LLMs for generating numerical predictions in robotics tasks, specifically for robotic grasping. We propose Reasoning Tuning, a novel method that integrates a reasoning phase before prediction during training, leveraging the extensive prior knowledge and advanced reasoning abilities of LLMs. This approach enables LLMs, notably with multi-modal capabilities, to generate accurate numerical outputs like grasp poses that are context-aware and adaptable through conversations. Additionally, we present the Reasoning Tuning VLM Grasp dataset, carefully curated to facilitate the adaptation of LLMs to robotic grasping. Extensive validation on both grasping datasets and real-world experiments underscores the adaptability of multi-modal LLMs for numerical prediction tasks in robotics. This not only expands their applicability but also bridges the gap between text-based planning and direct robot control, thereby maximizing the potential of LLMs in robotics. Our dataset will be released. More details and videos of this work are available on our project page: https://sites.google.com/view/rt-grasp. Jinxuan Xu, Shiyu Jin, Liangjun Zhang |
IROS | 2 |
| 2024 | Self-Supervised Learning for Graph Dataset CondensationabstractGraph dataset condensation (GDC) reduces a dataset with many graphs into a smaller dataset with fewer graphs while maintaining model training accuracy. GDC saves the storage cost and hence accelerates training. Although several GDC methods have been proposed, they are all supervised and require massive labels for the graphs, while graph labels can be scarce in many practical scenarios. To fill this gap, we propose a self-supervised graph dataset condensation method called SGDC, which does not require label information. Our initial design starts with the classical bilevel optimization paradigm for dataset condensation and incorporates contrastive learning techniques. But such a solution yields poor accuracy due to the biased gradient estimation caused by data augmentation. To solve this problem, we introduce representation matching, which conducts training by aligning the representations produced by the condensed graphs with the target representations generated by a pre-trained SSL model. This design eliminates the need for data augmentation and avoids biased gradient. We further propose a graph attention kernel, which not only improves accuracy but also reduces running time when combined with self-supervised kernel ridge regression (KRR). To simplify SGDC and make it more robust, we adopt a adjacency matrix reusing approach, which reuses the topology of the original graphs for the condensed graphs instead of repeatedly learning topology during training. Our evaluations on seven graph datasets find that SGDC improves model accuracy by up to 9.7% compared with 5 state-of-the-art baselines, even if they use label information. Moreover, SGDC is significantly more efficient than the baselines. Yuxiang Wang 0013, Xiao Yan 0002, Shiyu Jin, Hao Huang 0001, Quanqing Xu, Qingchen Zhang 0001, Bo Du 0001, Jiawei Jiang 0001 |
KDD | 3 |
| 2024 | The Integration of Macroscopic Traffic Optimization Control and Microscopic Traffic Flow Model for Mixed Traffic: A Cyber-Physical System PerspectiveabstractIn future transportation systems, there is a growing emphasis on integrating macroscopic traffic control and the state of microscopic traffic flow to improve traffic efficiency and safety. And cyber-physical system (CPS) can support integrated, efficient and reliable operations of diverse critical intelligent transportation infrastructure and services. Considering the limitations imposed by the communication range and computing capabilities of single agent, a novel method for combining macro traffic optimization with microscopic traffic flow models from a CPS perspective is presented. This approach takes into account not only exceptional events but also the propagation of congestion waves. Specifically, we analyze macroscopic traffic patterns and transmit optimization results to the microscopic level, as inputs to the microscopic traffic flow model, thereby influencing individual vehicle. Furthermore, the intelligent driver model, incorporating a molecular dynamics-based representation of how multiple leading vehicles impact the ego vehicle is investigated. By leveraging CPS-based macro traffic analysis, we improve the car-following model. To validate the effectiveness of the proposed optimization strategy, we conducted simulation experiments from both microscopic and macroscopic perspectives under mixed traffic environment. The results demonstrate the efficacy of the proposed approach in optimizing vehicle fuel consumption and improving traffic efficiency. Huamin Li, Shiyu Jin |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | GOATS: Goal Sampling Adaptation for Scooping with Curriculum Reinforcement LearningabstractIn this work, we first formulate the problem of robotic water scooping using goal-conditioned reinforcement learning. This task is particularly challenging due to the complex dynamics of fluid and the need to achieve multi-modal goals. The policy is required to successfully reach both position goals and water amount goals, which leads to a large convoluted goal state space. To overcome these challenges, we introduce Goal Sampling Adaptation for Scooping (GOATS), a curriculum reinforcement learning method that can learn an effective and generalizable policy for robot scooping tasks. Specifically, we use a goal-factorized reward formulation and interpolate position goal distributions and amount goal distributions to create curriculum throughout the learning process. As a result, our proposed method can outperform the baselines in simulation and achieves 5.46% and 8.71% amount errors on bowl scooping and bucket scooping tasks, respectively, under 1000 variations of initial water states in the tank and a large goal state space. Besides being effective in simulation environments, our method can efficiently adapt to noisy real-robot water-scooping scenarios with diverse physical configurations and unseen settings, demonstrating superior efficacy and generalizability. The videos of this work are available on our project page: https://sites.google.com/view/goatscooping. Yaru Niu, Shiyu Jin, Zeqing Zhang, Ding Zhao, Liangjun Zhang |
IROS | 2 |
| 2023 | CNN-Siam: multimodal siamese CNN-based deep learning approach for drug‒drug interaction predictionabstractBACKGROUND: Drug‒drug interactions (DDIs) are reactions between two or more drugs, i.e., possible situations that occur when two or more drugs are used simultaneously. DDIs act as an important link in both drug development and clinical treatment. Since it is not possible to study the interactions of such a large number of drugs using experimental means, a computer-based deep learning solution is always worth investigating. We propose a deep learning-based model that uses twin convolutional neural networks to learn representations from multimodal drug data and to make predictions about the possible types of drug effects. RESULTS: In this paper, we propose a novel convolutional neural network algorithm using a Siamese network architecture called CNN-Siam. CNN-Siam uses a convolutional neural network (CNN) as a backbone network in the form of a twin network architecture to learn the feature representation of drug pairs from multimodal data of drugs (including chemical substructures, targets and enzymes). Moreover, this network is used to predict the types of drug interactions with the best optimization algorithms available (RAdam and LookAhead). The experimental data show that the CNN-Siam achieves an area under the precision-recall (AUPR) curve score of 0.96 on the benchmark dataset and a correct rate of 92%. These results are significant improvements compared to the state-of-the-art method (from 86 to 92%) and demonstrate the robustness of the CNN-Siam and the superiority of the new optimization algorithm through ablation experiments. CONCLUSION: The experimental results show that our multimodal siamese convolutional neural network can accurately predict DDIs, and the Siamese network architecture is able to learn the feature representation of drug pairs better than individual networks. CNN-Siam outperforms other state-of-the-art algorithms with the combination of data enhancement and better optimizers. But at the same time, CNN-Siam has some drawbacks, longer training time, generalization needs to be improved, and poorer classification results on some classes. Kuiyuan Tong, Shiyu Jin, Shiyan Wang |
BMC Bioinform. | 3 |
| 2022 | Learning Insertion Primitives with Discrete-Continuous Hybrid Action Space for Robotic Assembly TasksabstractThis paper introduces a discrete-continuous action space to learn insertion primitives for robotic assembly tasks. Primitives are sequences of elementary actions with certain exit conditions, such as “pushing down the peg until contact”. Since the primitive is an abstraction of robot control commands and encodes human prior knowledge, it reduces the exploration difficulty and yields better learning efficiency. In this paper, we learn robot assembly skills via primitives. Specifically, we formulate insertion primitives as parameterized actions: hybrid actions consisting of discrete primitive types and continuous primitive parameters. Compared with the previous work using a set of discretized parameters for each primitive, the agent in our method can freely choose primitive parameters from a continuous space, which is more flexible and efficient. To learn these insertion primitives, we propose Twin-Smoothed Multi-pass Deep Q-Network (TS-MP-DQN), an advanced version of MP-DQN with twin Q-network to reduce the Q-value over-estimation. Extensive experiments are conducted in the simulation and real world for validation. From experiment results, our approach achieves higher success rates than three baselines: MP-DQN with parameterized actions, primitives with discrete parameters, and continuous velocity control. Furthermore, learned primitives are robust to sim-to-real transfer and can generalize to challenging assembly tasks such as tight round peg-hole and complex shaped electric connectors with promising success rates. Experiment videos are available at https://msc.berkeley.edu/research/insertion-primitives.html. Xiang Zhang 0020, Shiyu Jin, Xinghao Zhu, Masayoshi Tomizuka |
ICRA | 2 |
| 2022 | LCU-level Rate-Distortion Optimization for Versatile Video CodingabstractRate-distortion curve is greatly influenced by the texture and motion compensation. Therefore, using a uniform Lagrangian multiplier (lambda, λ to all the LCUs in a picture may not achieve the optimal coding performance. To address the challenging problem on how to formulate the appropriate λ to an individual LCU, this paper proposes a LCU-level Lagrangian multiplier adaption (LLMA) algorithm. Firstly, a motion compensation information entropy (MCIE) of 16 × 16 block is calculated by obtaining a simple binary information entropy from the source. Secondly, MCIE is normalized to a standard normal distribution, on the basis of the strong assumption that the video sources obey the Gaussian distribution. Thirdly, an adaption factor is assigned to a LCU, which multiples on the picture-level λ from the system calculation. Experiments show that the proposed LLMA algorithm achieves remarkable performance in low-delay and random-access configurations, with the resulting achievements of 1.81% BD-Rate gains, compared with VTM 13.0 at low-delay common test condition. Gencheng Xu, Shiyu Jin, Kaichen Tang, Yimin Zhou 0002 |
ISCAS | 2 |
| 2021 | Trajectory Optimization for Manipulation of Deformable Objects: Assembly of Belt Drive UnitsabstractThis paper presents a novel trajectory optimization formulation to solve the robotic assembly of the belt drive unit. Robotic manipulations involving contacts and deformable objects are challenging in both dynamic modeling and trajectory planning. For modeling, variations in the belt tension and contact forces between the belt and the pulley could dramatically change the system dynamics. For trajectory planning, it is computationally expensive to plan trajectories for such hybrid dynamical systems as it usually requires planning for discrete modes separately. In this work, we formulate the belt drive unit assembly task as a trajectory optimization problem with complementarity constraints to avoid explicitly imposing contact mode sequences. The problem is solved as a mathematical program with complementarity constraints (MPCC) to obtain feasible and efficient assembly trajectories. We validate the proposed method both in simulations with a physics engine and in real-world experiments with a robotic manipulator. Shiyu Jin, Diego Romeres, Arvind Ragunathan, Devesh K. Jha, Masayoshi Tomizuka |
ICRA | 1 |
| 2019 | Robust Deformation Model Approximation for Robotic Cable ManipulationabstractCable manipulation is a challenging task for robots. The major challenge is that cables have high degrees of freedom and are easy to deform during manipulation. In this paper, we propose a novel framework SPR-RWLS to manipulate cables, which includes real-time cable tracking and robust local deformation model approximation. For cable tracking, structure preserved registration (SPR) is utilized to robustly estimate the movement of selected points on a cable even in the presence of sensor noise, outliers, and occlusions. Robust weighted least squares (RWLS) is then applied to calculate the local deformation model of the cable under uncertainties. We show that SPR-RWLS enables the dual-arm robots to manipulate cables with different thicknesses and lengths to different desired curvatures in multiple scenarios. We also show that real-time implementation of the proposed method can be simplified by parallel computation. Shiyu Jin, Masayoshi Tomizuka |
IROS | 1 |