VLDB 2026 Research / reviewers in the wild / expert
Siyuan Dong
dblp:150/1489
· DBLP profile ↗
32ranked-venue papers
9as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 7 since 2021Systems, architecture and hardware · 12 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration
Siyuan Dong, Haotian Gong, Donna Pham |
Proc. VLDB Endow. | 2 |
| 2025 | Style mixup enhanced disentanglement learning for unsupervised domain adaptation in medical image segmentation
Zhuotong Cai, Jingmin Xin, Chenyu You, Peiwen Shi, Siyuan Dong, Nicha C. Dvornek, Nanning Zheng 0001, James S. Duncan |
Medical Image Anal. | 5 |
| 2025 | A Flow-based Truncated Denoising Diffusion Model for super-resolution Magnetic Resonance Spectroscopic ImagingabstractMagnetic Resonance Spectroscopic Imaging (MRSI) is a non-invasive imaging technique for studying metabolism and has become a crucial tool for understanding neurological diseases , cancers and diabetes. High spatial resolution MRSI is needed to characterize lesions, but in practice MRSI is acquired at low resolution due to time and sensitivity restrictions caused by the low metabolite concentrations. Therefore, there is an imperative need for a post-processing approach to generate high-resolution MRSI from low-resolution data that can be acquired fast and with high sensitivity. Deep learning-based super-resolution methods provided promising results for improving the spatial resolution of MRSI, but they still have limited capability to generate accurate and high-quality images. Recently, diffusion models have demonstrated superior learning capability than other generative models in various tasks, but sampling from diffusion models requires iterating through a large number of diffusion steps, which is time-consuming. This work introduces a Flow-based Truncated Denoising Diffusion Model (FTDDM) for super-resolution MRSI, which shortens the diffusion process by truncating the diffusion chain, and the truncated steps are estimated using a normalizing flow-based network. The network is conditioned on upscaling factors to enable multi-scale super-resolution. To train and evaluate the deep learning models, we developed a 1 H-MRSI dataset acquired from 25 high-grade glioma patients. We demonstrate that FTDDM outperforms existing generative models while speeding up the sampling process by over 9-fold compared to the baseline diffusion model. Neuroradiologists’ evaluations confirmed the clinical advantages of our method, which also supports uncertainty estimation and sharpness adjustment, extending its potential clinical applications. Siyuan Dong, Zhuotong Cai, Gilbert Hangel, Wolfgang Bogner, Georg Widhalm, Yaqing Huang, Qinghao Liang, Chenyu You, Chathura Kumaragamage, Robert K. Fulbright, Amit Mahajan, Amin Karbasi, John A. Onofrey, Robin A. de Graaf, James S. Duncan |
Medical Image Anal. | 1 |
| 2025 | PTDA: Progressive Pseudo-Label Learning for Cross-Domain Cloud Detection in High-Resolution Remote SensingabstractThe global cloud detection of high-resolution remote sensing images (HRSI) is crucial for acquiring high-quality imagery and optimizing data utilization. Traditional cloud detection models, which rely on limited samples and fully supervised learning, struggle to adapt to cross-temporal and cross-spatial domains. While current unsupervised domain adaptation (UDA) methods improve performance in cross-domain cloud detection to some extent, generating high-quality, reliable pseudo-labels remains a significant challenge for global cloud detection. Therefore, this paper proposes a progressive pseudo-label learning for cross-domain cloud detection in high-resolution remote sensing (PTDA). Firstly, we propose an online domain-invariant feature guided pseudo-label generation (OPLG) strategy and learning intra-domain unaligned features (LIUF), which effectively integrate domain-invariant features and intra-domain semantics to generate high-quality pseudo-labels at the feature level. LIUF then refines the pseudo-label quality at the pixel level. Secondly, during the model training, pseudo-label constrained intra-domain feature mining loss(PCIF Loss) is designed to suppress noisy semantic information within the domain, the hole effect of thick/thin clouds, and the noise interference of the contour boundary. Four cloud detection datasets, including MS Cloud (MS), HRC WHU Cloud (WHU), 95 Cloud(95), and WHUS2-CD+(S2), are grouped into three cross-domain tests, MS2WHU, MS2S2, and WHU295. Our approach achieved the best performance with mIoU 63.99%, 58.14%, 58.82%, and OA 80.03%, 79.36%, 80.49%, respectively. The experimental results show that the proposed method outperforms seven state-of-art cross-domain comparison methods. Thus, our method has important application value for cross-domain cloud detection. The available code can be downloaded from https://github.com/gasking/PTDA. Jin Kuang, Xianjun Gao, Yuanwei Yang, Siyuan Dong, Ji Dong, Yuan Kou, Meilin Tan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Per-Flow Quantile Estimation Using M4 FrameworkabstractThis paper introduces a novel framework, M4, designed to estimate per-flow quantiles in data streams accurately. M4 is a versatile framework that can be integrated with a wide array of single-flow quantile estimation algorithms, thereby enabling them to perform per-flow estimation. The framework employs a sketch-based approach to provide a space-efficient method for recording and extracting distribution information. M4 incorporates two techniques:MINIMUMandSUM. TheMINIMUMtechnique minimizes the noise on a flow from other flows caused by hash collisions, while theSUMtechnique efficiently categorizes flows based on their sizes and customizes treatment strategies accordingly. We demonstrate the application of M4 on three single-flow quantile estimation algorithms (DDSketch,$t$-digest, and ReqSketch), detailing the specific implementation of theMINIMUMandSUMtechniques. We provide theoretical proof that M4 delivers high accuracy while utilizing limited memory. Additionally, we conduct extensive experiments to evaluate the performance of M4 regarding accuracy and speed. The experimental results indicate that across all three example algorithms, M4 significantly outperforms two comparison frameworks in terms of accuracy for per-flow quantile estimation while maintaining comparable speed. Zhuochen Fan, Yalun Cai, Siyuan Dong, Qiuheng Yin, Tianyu Bai, Hanyu Xue, Peiqing Chen, Yuhan Wu 0001, Tong Yang 0003, Bin Cui 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Symmetric Consistency with Cross-Domain Mixup for Cross-Modality Cardiac SegmentationabstractAccurate cardiac segmentation in cross-modality images plays an important role in the quantitative analysis of the heart to diagnose cardiovascular diseases. However, achieving high performance in cross-modality segmentation is hindered by the time-consuming annotation and modality gap. While some approaches employ Unsupervised Domain Adaptation (UDA) through adversarial learning to address the issue, it still remains challenging due to the instability of the adversarial generative models. In this work, we propose Symmetric Consistency with Cross-Domain Mixup (SCCDM), integrated with the teacher-student model for cross-modality cardiac segmentation. Specifically, we introduce symmetric consistency across the domains for two mixed data to diversify the data distribution from both the source domain and target domain. Extensive experiments on a public cardiac dataset demonstrate that SCCDM achieves superior domain adaptation performance for cardiac segmentation compared to state-of-the-art methods. Zhuotong Cai, Jingmin Xin, Siyuan Dong, John A. Onofrey, Nanning Zheng 0001, James S. Duncan |
ICASSP | 3 |
| 2024 | M4: A Framework for Per-Flow Quantile EstimationabstractThe field of quantile estimation has grown in importance due to its myriad practical applications. Recent research trends have evolved from estimating the quantile for a single data stream to developing data structures that can concurrently estimate quantiles for multiple sub-streams, also known as flows. This paper introduces a novel framework, M4, designed to estimate per-flow quantiles in data streams accurately. M4 is a versatile framework that can be integrated with a wide array of single-flow quantile estimation algorithms, thereby enabling them to perform per-flow estimation. The framework employs a sketch-based approach to provide a space-efficient method for recording and extracting distribution information. M4 incorporates two techniques: MINIMUM and SUM. The MINIMUM technique minimizes the noise on a flow from other flows caused by hash collisions, while the SUM technique efficiently categorizes flows based on their sizes and customizes treatment strategies accordingly. We demonstrate the application of M4 on three single-flow quantile estimation algorithms (DDSketch, t-digest, and ReqSketch), detailing the specific implementation of the MINIMUM and SUM techniques. We provide theoretical proof that M4 delivers high accuracy while utilizing limited memory. Additionally, we conduct extensive experiments to evaluate the performance of M4 regarding accuracy and speed. The experimental results indicate that across all three example algorithms, M4 significantly outperforms two comparison frameworks in terms of accuracy for per-flow quantile estimation while maintaining comparable speed. Siyuan Dong, Zhuochen Fan, Tianyu Bai, Tong Yang 0003, Hanyu Xue, Peiqing Chen, Yuhan Wu 0001 |
ICDE | 1 |
| 2024 | Class-Aware Mutual Mixup with Triple Alignments for Semi-supervised Cross-Domain Segmentation
Zhuotong Cai, Jingmin Xin, Tianyi Zeng, Siyuan Dong, Nanning Zheng 0001, James S. Duncan |
MICCAI (8) | 4 |
| 2024 | 3-D Gravity and Magnetic Joint Inversion Based on Deep Learning Combined With Measurement Data ConstraintabstractThe joint inversion of gravity and magnetic data can reduce the nonuniqueness problem of potential field data inversion. We propose a gravity and magnetic joint inversion method based on deep learning (DL) combined with measurement data constraint. The framework obtains the gravity and magnetic dataset required for network training by randomly generating the underground structural consistency model and then inputs the dataset into the network for training. Moreover, we add constraints to the measurement data in the training of the network, that is, fitting the data anomalies obtained by the inversion model through forward calculation with the real anomalies, which makes the network more consistent with geophysical theory. In the test phase, the trained network can obtain the inversion results rapidly, and the inversion results of the testing dataset show that this method can obtain better results when applied to the joint inversion of gravity and magnetic fields than the conventional regularized inversion and cross-gradient joint inversion methods. In addition, our method can also distinguish the anomaly conditions in the case of structural inconsistency. Furthermore, we apply this method to actual gravity and magnetic data of Gonghe Basin, Qinghai Province, China, and predict the distribution of dry hot rock related to geothermal resources. Siyuan Dong, Zhaofa Zeng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Magnetic Data Edge Detection Method With Depth Information Based on UNet ++abstractEdge detection is a critical technology in processing potential field data, enabling the rapid identification of geological body edges using magnetic anomaly data. Traditional methods for detecting edges in magnetic data are known for their simplicity and efficiency; however, they suffer from low resolution, poor robustness, and a lack of depth information. In recent years, the application of deep learning (DL) to edge detection in field data has enhanced both the resolution and robustness of these methods. Nonetheless, these approaches still fail to determine the buried depths of geological body edges. To address this issue, this study has developed an edge detection method called multiconstraint DL based on UNet++, which not only identifies the edge positions but also ascertains the buried depths of geological bodies. The study proposes the multiconstraint loss function for edge detection and employs the Dice loss function and the mean square error (mse) loss function to jointly supervise the network’s training. Subsequently, the label design was revised to include depth information on the geological body, enabling the DL method to accurately determine the buried depths of geological bodies. Analysis of the test model’s detection results reveals that the edge detection method based on UNet++ can precisely identify both the edge positions and the exact buried depths of geological bodies, which makes up for the shortcoming that the traditional edge detection method does not contain the depth information. Moreover, this method resolves the issue of discontinuity found in traditional DL edge detection. Finally, the method was applied to actual aeromagnetic data from the Dandong area in Liaoning Province, successfully identifying the edge positions of the main mining area and providing targeted regions for further exploration. Xiuan Yao, Xiangcheng Zeng, Siyuan Dong, Zhenyu Yu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Unsupervised Domain Adaptation by Cross-Prototype Contrastive Learning for Medical Image SegmentationabstractUnsupervised Domain Adaptation (UDA), which aligns the labeled source distribution to the unlabeled target distribution, has shown remarkable achievement in the medical image segmentation task. Previous UDA methods unilaterally consider the global distribution alignment through explicit category-based loss while good separation and discrimination of class are insufficiently explored, resulting in the sub-aligned distribution across domains. In this paper, we propose cross-prototype contrastive learning method (CPCL) for UDA segmentation through class centroid alignment. Specifically, to reduce the intra-class distance and increase the inter-class distance, we first introduce prototype-feature contrastive learning to align the pixel-level features and the same-class global prototype across domains. Secondly, we further present prototype-prototype contrastive learning to align the same class prototypes between the source domain and target domain for compact category centroid and better global domain distribution alignment. Extensive experiments on two public cardiac datasets demonstrate that the proposed CPCL achieves superior domain adaptation performance as compared with the state-of-the-art. Zhuotong Cai, Jingmin Xin, Siyuan Dong, Chenyu You, Peiwen Shi, Tianyi Zeng, John A. Onofrey, Nanning Zheng 0001, James S. Duncan |
BIBM | 3 |
| 2023 | KVSAgg: Secure Aggregation of Distributed Key-Value SetsabstractIn global data analysis, the central server needs the global statistic of the user data stored in local clients. In such cases, an Honest-but-Curious central server might put user privacy at risk in trying to collect individual statistics of each user. In response, the secure aggregation provides a solution for calculating global statistics without revealing users’ privacy data. However, existing secure aggregation protocols only focus on the data in the form of vectors or common sets, which limits their application scope. We formalize a general problem—key-value set secure aggregation—that not only includes secure vector aggregation and private set union but also supports more applications. To address the proposed problem, we devise our solution (called the KVSAgg framework) that promises satisfactory performance in security, efficiency, and accuracy. Our key technique is a homomorphic transform algorithm (called HyperIBLT) that is not only capable of bidirectionally transforming data between key-value sets and vectors, but also able to transform sum operation of sets to addition of vectors. We implement KVSAgg on both CPU and GPU platforms and perform the evaluation on three use cases including federated learning, distributed data counting, and finding global hot items. Compared with our baselines, KVSAgg simultaneously achieves the best security, efficiency higher by orders of magnitude, and zero-error in nearly all cases. All codes are open-source anonymously. Yuhan Wu 0001, Siyuan Dong, Yikai Zhao 0001, Fangcheng Fu, Tong Yang 0003, Chaoyue Niu, Fan Wu 0006, Bin Cui 0001 |
ICDE | 2 |
| 2023 | Neural Contact Fields: Tracking Extrinsic Contact with Tactile SensingabstractWe present Neural Contact Fields, a method that brings together neural fields and tactile sensing to address the problem of tracking extrinsic contact between object and environment. Knowing where the external contact occurs is a first step towards methods that can actively control it in facilitating downstream manipulation tasks. Prior work for localizing environmental contacts typically assume a contact type (e.g. point or line), does not capture contact/no-contact transitions, and only works with basic geometric-shaped objects. Neural Contact Fields are the first method that can track arbitrary multi-modal extrinsic contacts without making any assumptions about the contact type. Our key insight is to estimate the probability of contact for any 3D point in the latent space of object's shapes, given vision-based tactile inputs that sense the local motion resulting from the external contact. In experiments, we find that Neural Contact Fields are able to localize multiple contact patches without making any assumptions about the geometry of the contact, and capture contact/no-contact transitions for known categories of objects with unseen shapes in unseen environment configurations. In addition to Neural Contact Fields, we also release our YCB-Extrinsic-Contact dataset of simulated extrinsic contact interactions to enable further research in this area. Project page: https://github.com/carolinahiguera/NCF Carolina Higuera, Siyuan Dong, Byron Boots, Mustafa Mukadam |
ICRA | 2 |
| 2023 | MicroscopeSketch: Accurate Sliding Estimation Using Adaptive ZoomingabstractHigh-accuracy real-time data stream estimations are critical for various applications, and sliding-window-based techniques have attracted wide attention. However, existing solutions struggle to achieve high accuracy, generality, and low memory usage simultaneously. To overcome these limitations, we present MicroscopeSketch, a high-accuracy sketch framework. Our key technique, called adaptive zooming, dynamically adjusts the granularity of counters to maximize accuracy while minimizing memory usage. By applying MicroscopeSketch to three specific tasks---frequency estimation, top-k frequent items discovery, and top-k heavy changes identification-we demonstrate substantial improvements over existing methods, reducing errors by roughly 4 times for frequency estimation and 3 times for identifying top-k items. The relevant source code is available in a GitHub repository. Yuhan Wu 0001, Shiqi Jiang 0004, Siyuan Dong, Jiale Chen 0003, Yutong Hu 0002, Tong Yang 0003, Steve Uhlig, Bin Cui 0001 |
KDD | 3 |
| 2023 | 3-D Gravity Data Inversion Based on Enhanced Dual U-Net FrameworkabstractThree-dimensional gravity inversion is an effective method for restoring underground density distribution from gravity anomaly data. Conventional regularization inversion has good data fitting, but its inversion model has insufficient model fitting capabilities due to its low-depth resolution. Although data-driven deep learning-based gravity inversion results significantly improve depth resolution and physical property distribution, it is difficult to ensure the data fitting of the inversion results. Accordingly, this study proposes a three-dimensional gravity data inversion based on enhanced dual U-Net framework (EdU-Net) to solve the above problems, making the inversion results have good model and data fitting performance. The proposed EdU-Net consists of two parts: first, training a large generalization pre-trained network Net I, and then quickly generating an enhanced Net II for the target data through fine-tuning. Additionally, this study adds forward-fitting constraints in the framework’s loss function to reduce the problem of large data-fitting errors in traditional data-driven deep learning inversion. The trained Net II inversion result has better model and data fitting accuracy than Net I. Moreover, by comparing the inversion results of synthetic models, this study demonstrates that the EdU-Net method performs better than traditional deep learning. Finally, this method is applied to the measured data of the Gonghe Basin in Qinghai Province, China, and provides a reasonable explanation for the distribution of hot dry rocks. Siyuan Dong, Pengyu Lu, Zhaofa Zeng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | HoppingSketch: More Accurate Temporal Membership Query and Frequency QueryabstractNowadays, research on temporal membership queries is indispensable. Generally, temporal membership queries exist in two modalities: fixed windows and sliding windows, the latter having obvious advantages. The first sketch that implements temporal membership queries is the persistent Bloom filter (PBF). PBF has two shortcomings: it does not support sliding windows nor frequency queries. Here, we propose HoppingSketch to promote the original PBF. It is the first sketch that implements temporal membership queries for sliding windows. HoppingSketch is a general and efficient data stream processing framework, able to implement different tasks thanks to different atomic sketches. When the atomic sketches are Bloom filters and we apply them to PBF, HoppingSketch can achieve significantly higher temporal membership query accuracy than the original PBF. When the atomic sketches are sketches of Count-Min, Conservative Update, and Count, HoppingSketch can achieve more accurate frequency query than by applying PBF on the corresponding sketches. Our experimental results demonstrate the advantages of HoppingSketch compared with the state-of-the-art. Zhuochen Fan, Siyuan Dong, Fangyi Liu, Tong Yang 0003, Steve Uhlig, Bin Cui 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | GelSlim 3.0: High-Resolution Measurement of Shape, Force and Slip in a Compact Tactile-Sensing FingerabstractThis work presents a new version of tactile-sensing finger, GelSlim 3.0, which integrates the ability to sense high-resolution shape, force, and slip in a more compact form factor than previous implementations, designed for cluttered bin-picking scenarios. The novel design integrates real-time model-based algorithms to measure shape, estimate the 3-D contact force distribution, and detect incipient slip. The constraints imposed by the photometric stereo algorithm used for depth reconstruction and the implementation of a planar sensing surface make the miniaturization of previous designs nontrivial. To achieve a compact integration, we optimize the optical path from illumination source to camera. Using an optical simulation environment, we develop an illumination shaping lens and position the source LEDs and camera. The optimized optical configuration is integrated into a finger design composed of a robust and easily replaceable snap-to-fit fingertip module that facilitates manufacture, assembly, use, and repair. To stimulate future research in tactile-sensing and provide the robotics community access to a reliable and easily reproducible tactile finger with a diversity of sensing modalities, we open-source the design, fabrication methods, and software at https://github.com/mcubelab/gelslim. Ian H. Taylor, Siyuan Dong, Alberto Rodriguez 0003 |
ICRA | 2 |
| 2022 | Visual-Tactile Multimodality for Following Deformable Linear Objects Using Reinforcement LearningabstractManipulation of deformable objects is a challenging task for a robot. It would be problematic to use a single sensory input to track the behaviour of such objects: vision can be subjected to occlusions, whereas tactile inputs cannot capture the global information that is useful for the task. In this paper, we study the problem of using vision and tactile inputs together to complete the task of following deformable linear objects, for the first time. We create a Reinforcement Learning agent using different sensing modalities and investigate how its behaviour can be boosted using visual-tactile fusion, compared to using a single sensing modality. To this end, we developed a benchmark in simulation for manipulating the deformable linear objects using multimodal sensing inputs. The policy of the agent uses distilled information, e.g., the pose of the object in both visual and tactile perspectives, instead of the raw sensing signals, so that it can be directly transferred to real environments. In this way, we disentangle the perception system and the learned control policy. Our extensive experiments show that the use of both vision and tactile inputs, together with proprioception, allows the agent to complete the task in up to 92% of cases, compared to 77% when only one of the signals is given. Our results can provide valuable insights for the future design of tactile sensors and for deformable objects manipulation. Code and videos can be found at: https://github.com/lpecyna/SoftSlidingGym. Leszek Pecyna, Siyuan Dong, Shan Luo 0001 |
IROS | 2 |
| 2022 | Invertible Sharpening Network for MRI Reconstruction Enhancement
Siyuan Dong, Eric Z. Chen, Lin Zhao 0004, Xiao Chen 0013, Yikang Liu 0001, Terrence Chen, Shanhui Sun |
MICCAI (6) | 1 |
| 2022 | Multi-scale Super-Resolution Magnetic Resonance Spectroscopic Imaging with Adjustable Sharpness
Siyuan Dong, Gilbert Hangel, Wolfgang Bogner, Georg Widhalm, Karl Rössler, Siegfried Trattnig, Chenyu You, Robin A. de Graaf, John A. Onofrey, James S. Duncan |
MICCAI (6) | 1 |
| 2022 | Class-Aware Adversarial Transformers for Medical Image SegmentationabstractTransformers have made remarkable progress towards modeling long-range dependencies within the medical image analysis domain. However, current transformer-based models suffer from several disadvantages: (1) existing methods fail to capture the important features of the images due to the naive tokenization scheme; (2) the models suffer from information loss because they only consider single-scale feature representations; and (3) the segmentation label maps generated by the models are not accurate enough without considering rich semantic contexts and anatomical textures. In this work, we present CASTformer, a novel type of adversarial transformers, for 2D medical image segmentation. First, we take advantage of the pyramid structure to construct multi-scale representations and handle multi-scale variations. We then design a novel class-aware transformer module to better learn the discriminative regions of objects with semantic structures. Lastly, we utilize an adversarial training strategy that boosts segmentation accuracy and correspondingly allows a transformer-based discriminator to capture high-level semantically correlated contents and low-level anatomical features. Our experiments demonstrate that CASTformer dramatically outperforms previous state-of-the-art transformer-based approaches on three benchmarks, obtaining 2.54%-5.88% absolute improvements in Dice over previous models. Further qualitative experiments provide a more detailed picture of the model’s inner workings, shed light on the challenges in improved transparency, and demonstrate that transfer learning can greatly improve performance and reduce the size of medical image datasets in training, making CASTformer a strong starting point for downstream medical image analysis tasks. Chenyu You, Ruihan Zhao 0001, Siyuan Dong, Sandeep Chinchali, Ufuk Topcu, Lawrence H. Staib, James S. Duncan |
NeurIPS | 4 |
| 2021 | Tactile-RL for Insertion: Generalization to Objects of Unknown GeometryabstractObject insertion is a classic contact-rich manipulation task. The task remains challenging, especially when considering general objects of unknown geometry, which significantly limits the ability to understand the contact configuration between the object and the environment. We study the problem of aligning the object and environment with a tactile-based feedback insertion policy. The insertion process is modeled as an episodic policy that iterates between insertion attempts followed by pose corrections. We explore different mechanisms to learn such a policy based on Reinforcement Learning. The key contribution of this paper is to demonstrate that it is possible to learn a tactile insertion policy that generalizes across different object geometries, and an ablation study of the key design choices for the learning agent: 1) the type of learning scheme: supervised vs. reinforcement learning; 2) the type of learning schedule: unguided vs. curriculum learning ; 3) the type of sensing modality: force/torque vs. tactile; and 4) the type of tactile representation: tactile RGB vs. tactile flow. We show that the optimal configuration of the learning agent (RL + curriculum + tactile flow) exposed to 4 training objects yields an closed-loop insertion policy that inserts 4 novel objects with over 85.0% success rate and within 3~4 consecutive attempts. Comparisons between F/T and tactile sensing, shows that while an F/T-based policy learns more efficiently, a tactile-based policy provides better generalization. See supplementary video and results at https://sites.google.com/view/tactileinsertion. Siyuan Dong, Devesh K. Jha, Diego Romeres, Sangwoon Kim, Daniel Nikovski, Alberto Rodriguez 0003 |
ICRA | 1 |
| 2021 | Extrinsic Contact Sensing with Relative-Motion Tracking from Distributed Tactile MeasurementsabstractThis paper addresses the localization of contacts of an unknown grasped rigid object with its environment, i.e., extrinsic to the robot. We explore the key role that distributed tactile sensing plays in localizing contacts external to the robot, in contrast to the role that aggregated force/torque measurements traditionally play in localizing contacts on the robot. When in contact with the environment, an object will move in accordance with the kinematic and possibly frictional constraints imposed by that contact. Small motions of the object, which are observable with tactile sensors, indirectly encode those constraints and the geometry that defines them.We formulate the extrinsic contact sensing problem as a constraint-based estimation problem. The estimation is subject to the kinematic constraints imposed by the tactile measurements of object motion, as well as the kinematic (e.g., non-penetration) and possibly frictional (e.g., sticking) constraints imposed by rigid-body mechanics. We validate the approach in simulation and with real experiments on the case studies of fixed point and line contacts.This paper discusses the theoretical basis for the value of distributed tactile sensing in contrast to aggregated force/torque measurements. It also provides an estimation framework for localizing environmental contacts with potential impact in contact-rich manipulation scenarios such as assembling or packing. Daolin Ma, Siyuan Dong, Alberto Rodriguez 0003 |
ICRA | 2 |
| 2020 | Tactile Dexterity: Manipulation Primitives with Tactile FeedbackabstractThis paper develops closed-loop tactile controllers for dexterous robotic manipulation with a dual-palm robotic system. Tactile dexterity is an approach to dexterous manipulation that plans for robot/object interactions that render interpretable tactile information for control. We divide the role of tactile control into two goals: 1) control the contact state between the end-effector and the object (contact/no-contact, stick/slip) by regulating the stability of planned contact configurations and monitoring undesired slip events; and 2) control the object state by tactile-based tracking and iterative replanning of the object and robot trajectories. Key to this formulation is the decomposition of manipulation plans into sequences of manipulation primitives with simple mechanics and efficient planners. We consider the scenario of manipulating an object from an initial pose to a target pose on a flat surface while correcting for external perturbations and uncertainty in the initial pose of the object. We experimentally validate the approach with an ABB YuMi dual-arm robot and demonstrate the ability of the tactile controller to react to external perturbations. Francois Robert Hogan, José Ballester, Siyuan Dong, Alberto Rodriguez 0003 |
ICRA | 3 |
| 2019 | Maintaining Grasps within Slipping Bounds by Monitoring Incipient SlipabstractIn this paper, we propose an approach to detect incipient slip, i.e. predict slip, by using a high-resolution vision-based tactile sensor, GelSlim. The sensor dynamically captures the tactile imprints of the grasped object and their changes with a soft gel pad. The method assumes the object is mostly rigid and expects the motion of object's imprint on the sensor surface to be a 2D rigid-body motion. We use the deviation of the true motion field from that of a 2D planar rigid transformation as a measure of slip. The output is a dense slip field which we monitor in real time to detect when small areas of the contact patch start to slip (incipient slip). The method can detect incipient slip in any direction without any prior knowledge of the object at 24 Hz. We test the method on 10 objects for 240 times and achieve 86.25% detection accuracy with the vast majority of failure cases occurring when grasping highly deformable objects. We further show how the slip feedback can be used to adjust the gripping force to avoid slip with a closed-loop bottle-cap screwing and unscrewing experiment. The method can be used to enable many manipulation tasks in both structured and unstructured environments. Siyuan Dong, Daolin Ma, Elliott Donlon, Alberto Rodriguez 0003 |
ICRA | 1 |
| 2019 | Dense Tactile Force Estimation using GelSlim and inverse FEMabstractIn this paper, we present a new version of tactile sensor GelSlim 2.0 with the capability to estimate the contact force distribution in real time. The sensor is vision-based and uses an array of markers to track deformations on a gel pad due to contact. A new hardware design makes the sensor more rugged, parametrically adjusTable AND Improves illumination. leveraging the sensor's increased functionality, we propose to use inverse finite element method (ifem), a numerical method to reconstruct the contact force distribution based on marker displacements. the sensor is able to provide force distribution of contact with high spatial density. experiments and comparison with ground truth show that the reconstructed force distribution is physically reasonable with good accuracy.A sequence of Kendama manipulations with corresponding displacement field (yellow) and force field (red). Video can be found on Youtube: https://youtu.be/hWw9A0ZBZuU. Daolin Ma, Elliott Donlon, Siyuan Dong, Alberto Rodriguez 0003 |
ICRA | 3 |
| 2019 | Tactile-Based Insertion for Dense Box-PackingabstractWe study the problem of using high-resolution tactile sensors to control the insertion of objects in a boxpacking scenario. In this paper, we propose an insertion strategy that leverages tactile sensing to: 1) safely probe the box with the grasped object while monitoring incipient slip to maintain a stable grasp on the object. 2) estimate and correct for residual position uncertainties to insert the object into a designated gap without disturbing the environment.Our proposed methodology is based on two neural networks that estimate the error direction and error magnitude, from a stream of tactile imprints, acquired by two GelSlim fingers, during the insertion process. The system is trained on four objects with basic geometric shapes, which we show generalizes to four other common objects. Based on the estimated positional errors, a heuristic controller iteratively adjusts the position of the object and eventually inserts it successfully without requiring prior knowledge of the geometry of the object. The key insight is that dense tactile feedback contains useful information with respect to the contact interaction between the grasped object and its environment. We achieve high success rate and show that unknown objects can be inserted with an average of 6 attempts of the probe-correct loop. The method's ability to generalize to novel objects makes it a good fit for box packing in warehouse automation. Siyuan Dong, Alberto Rodriguez 0003 |
IROS | 1 |
| 2018 | Slip Detection with Combined Tactile and Visual InformationabstractSlip detection plays a vital role in robotic manipulation and it has long been a challenging problem in the robotic community. In this paper, we propose a new method based on deep neural network (DNN) to detect slip. The training data is acquired by a GelSight tactile sensor and a camera mounted on a gripper when we use a robot arm to grasp and lift 94 daily objects with different grasping forces and grasping positions. The DNN is trained to classify whether a slip occurred or not. To evaluate the performance of the DNN, we test 10 unseen objects in 152 grasps. A detection accuracy as high as 88.03 % is achieved. It is anticipated that the accuracy can be further improved with a larger dataset. This method is beneficial for robots to make stable grasps, which can be widely applied to automatic force control, grasping strategy selection and fine manipulation. Siyuan Dong, Edward H. Adelson |
ICRA | 2 |
| 2018 | GelSlim: A High-Resolution, Compact, Robust, and Calibrated Tactile-sensing FingerabstractThis work describes the development of a high-resolution tactile-sensing finger for robot grasping. This finger, inspired by previous GelSight sensing techniques (Johnson and Adelson 2009), features an integration that is slimmer, more robust, and with more homogeneous output than previous vision-based tactile sensors. To achieve a compact integration, we redesign the optical path from illumination source to camera by combining light guides and an arrangement of mirror reflections. We parameterize the optical path with geometric design variables and describe the tradeoffs between the finger thickness, camera depth of field, and size of the tactile sensing area. The sensor sustains the wear from continuous use - and abuse - in grasping tasks by combining tougher materials for the compliant gel, a textured fabric skin, a structurally rigid body, and a calibration process that maintains homogeneous illumination and contrast of the tactile images during use. Finally, we evaluate the sensor's durability along four metrics that track the signal quality during more than 3000 grasping experiments. Elliott Donlon, Siyuan Dong, Melody Liu, Edward H. Adelson, Alberto Rodriguez 0003 |
IROS | 2 |
| 2017 | Connecting Look and Feel: Associating the Visual and Tactile Properties of Physical MaterialsabstractFor machines to interact with the physical world, they must understand the physical properties of objects and materials they encounter. We use fabrics as an example of a deformable material with a rich set of mechanical properties. A thin flexible fabric, when draped, tends to look different from a heavy stiff fabric. It also feels different when touched. Using a collection of 118 fabric samples, we captured color and depth images of draped fabrics along with tactile data from a high-resolution touch sensor. We then sought to associate the information from vision and touch by jointly training CNNs across the three modalities. Through the CNN, each input, regardless of the modality, generates an embedding vector that records the fabrics physical property. By comparing the embedding vectors, our system is able to look at a fabric image and predict how it will feel, and vice versa. We also show that a system jointly trained on vision and touch data can outperform a similar system trained only on visual data when tested purely with visual inputs. Wenzhen Yuan 0001, Shaoxiong Wang, Siyuan Dong, Edward H. Adelson |
CVPR | 3 |
| 2017 | Improved GelSight tactile sensor for measuring geometry and slipabstractA GelSight sensor uses an elastomeric slab covered with a reflective membrane to measure tactile signals. It measures the 3D geometry and contact force information with high spacial resolution, and successfully helped many challenging robot tasks. A previous sensor [1], based on a semi-specular membrane, produces high resolution but with limited geometry accuracy. In this paper, we describe a new design of GelSight for robot gripper, using a Lambertian membrane and new illumination system, which gives greatly improved geometric accuracy while retaining the compact size. We demonstrate its use in measuring surface normals and reconstructing height maps using photometric stereo. We also use it for the task of slip detection, using a combination of information about relative motions on the membrane surface and the shear distortions. Using a robotic arm and a set of 37 everyday objects with varied properties, we find that the sensor can detect translational and rotational slip in general cases, and can be used to improve the stability of the grasp. Siyuan Dong, Wenzhen Yuan 0001, Edward H. Adelson |
IROS | 1 |
| 2014 | SEK: sparsity exploiting k-mer-based estimation of bacterial community compositionabstractMOTIVATION: Estimation of bacterial community composition from a high-throughput sequenced sample is an important task in metagenomics applications. As the sample sequence data typically harbors reads of variable lengths and different levels of biological and technical noise, accurate statistical analysis of such data is challenging. Currently popular estimation methods are typically time-consuming in a desktop computing environment. RESULTS: Using sparsity enforcing methods from the general sparse signal processing field (such as compressed sensing), we derive a solution to the community composition estimation problem by a simultaneous assignment of all sample reads to a pre-processed reference database. A general statistical model based on kernel density estimation techniques is introduced for the assignment task, and the model solution is obtained using convex optimization tools. Further, we design a greedy algorithm solution for a fast solution. Our approach offers a reasonably fast community composition estimation method, which is shown to be more robust to input data variation than a recently introduced related method. AVAILABILITY AND IMPLEMENTATION: A platform-independent Matlab implementation of the method is freely available at http://www.ee.kth.se/ctsoftware; source code that does not require access to Matlab is currently being tested and will be made available later through the above Web site. Saikat Chatterjee, David Koslicki, Siyuan Dong, Nicolas Innocenti, Lu Cheng 0004, Yueheng Lan, Mikko Vehkaperä, Mikael Skoglund, Lars K. Rasmussen, Erik Aurell, Jukka Corander |
Bioinform. | 3 |