EDBT 2026 Demo / reviewers in the wild / expert
Jieji Ren
dblp:278/8428
· DBLP profile ↗
16ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-6381-6830ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-Informed Token Prediction-Based Dynamic Modeling and High-Speed Feedforward Tracking Control of Dielectric Elastomer ActuatorsabstractDue to their continuous electromechanical deformation, rate-dependent viscoelasticity, and complex mechanical vibration, dynamic modeling and high-speed tracking control of dielectric elastomer actuators (DEAs) remain elusive, significantly limiting their working bandwidth. In this work, we propose a Physics-Informed Token Prediction (PITP) that enables accurate modeling of DEA dynamics and high-speed feedforward tracking control. The PITP framework consists of two key components: a physics-informed encoder and a dynamic decoder. The physics-informed encoder is designed based on a simplified equivalent linear model and trained through the hierarchical optimization training method, which embeds the global dynamic characteristics into tokens, minimizing the need for extensive data and training. Then, the dynamic decoder is developed by using these tokens as state-dependent parameters, capable of describing complex dynamic responses through the autoregressive solution. Finally, by taking advantage of the model's reversibility, a direct inverse compensator is established to linearize the input-output relationship. Experimental results of several DEAs with different configurations and payloads demonstrate that, based on our PITP framework, the complex nonlinear dynamic responses of all DEAs can be precisely described and eliminated within their natural frequency, validating its generality and versatility. By leveraging fast modeling ($< $30 minutes) and high-speed feedforward tracking control, our PITP framework may accelerate DEAs' practical applications. Xiaotian Shi, Peinan Yan, Jieji Ren, Guo-Ying Gu |
IEEE Trans. Robotics | 4 |
| 2025 | EventUPS: Uncalibrated Photometric Stereo Using an Event Camera
Jinxiu Liang, Bohan Yu, Haotian Zhuang, Jieji Ren, Peiqi Duan 0002, Boxin Shi |
ICCV | 5 |
| 2025 | Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial EnvironmentsabstractAnomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects in complex and unstructured industrial environments with varying views, poses and illumination remains challenging. We propose a novel anomaly detection and localization method specifically designed to handle inputs with perturbative patterns. Our approach introduces a new framework based on a collaborative distillation heterogeneous teacher network (HetNet), an adaptive local-global feature fusion module, and a local multivariate Gaussian noise generation module. HetNet can learn to model the complex feature distribution of normal patterns using limited information about local disruptive changes. We conducted extensive experiments on mainstream benchmarks. HetNet demonstrates superior performance with approximately 10% improvement across all evaluation metrics on MSC-AD under industrial conditions, while achieving state-of-the-art results on other datasets, validating its resilience to environmental fluctuations and its capability to enhance the reliability of industrial anomaly detection systems across diverse scenarios. Tests in real-world environments further confirm that HetNet can be effectively integrated into production lines to achieve robust and real-time anomaly detection. Codes, images and videos are published on the project website at: https://zihuatanejoyu.github.io/HetNet/ Jiawen Yu, Jieji Ren, Yang Chang, Qiaojun Yu, Xuan Tong, Boyang Wang 0003, Xinji Mai |
IROS | 2 |
| 2025 | ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationabstractVision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncertainty. To address these limitations, we propose \textbf{ForceVLA}, a novel end-to-end manipulation framework that treats external force sensing as a first-class modality within VLA systems. ForceVLA introduces \textbf{FVLMoE}, a force-aware Mixture-of-Experts fusion module that dynamically integrates pretrained visual-language embeddings with real-time 6-axis force feedback during action decoding. This enables context-aware routing across modality-specific experts, enhancing the robot's ability to adapt to subtle contact dynamics. We also introduce \textbf{ForceVLA-Data}, a new dataset comprising synchronized vision, proprioception, and force-torque signals across five contact-rich manipulation tasks. ForceVLA improves average task success by 23.2\% over strong $\pi_0$-based baselines, achieving up to 80\% success in tasks such as plug insertion. Our approach highlights the importance of multimodal integration for dexterous manipulation and sets a new benchmark for physically intelligent robotic control. Code and data will be released at https://sites.google.com/view/forcevla2025/. Jiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren, Ce Hao, Haitong Ding, Guangyu Huang, Guofan Huang, Panpan Cai, Cewu Lu |
NeurIPS | 4 |
| 2024 | DiLiGenRT: A Photometric Stereo Dataset with Quantified Roughness and TranslucencyabstractPhotometric stereo faces challenges from non-Lambertian reflectance in real-world scenarios. Systematically measuring the reliability of photometric stereo methods in handling such complex reflectance necessitates a real-world dataset with quantitatively controlled reflectances. This paper introduces DiLiGenRT, the first real-world dataset for evaluating photometric stereo methods under quantified reflectances by manufacturing 54 hemispheres with varying degrees of two reflectance properties: Roughness and Transluency, Unlike qualitative and semantic labels, such as “diffuse” and “specular,” that have been used in previous datasets, our quantified dataset allows comprehensive and systematic benchmark evaluations. In addition, it facilitates selecting best-fit photometric stereo methods based on the quantitative reflectance properties. Our dataset and benchmark results are available at https://photometricstereo.github.io/diligentrt.html. Heng Guo 0003, Jieji Ren, Feishi Wang, Boxin Shi, Ming Jun Ren, Yasuyuki Matsushita |
CVPR | 2 |
| 2024 | EventPS: Real-Time Photometric Stereo Using an Event CameraabstractPhotometric stereo is a well-established technique to es-timate the surface normal of an object. However, the re-quirement of capturing multiple high dynamic range images under different illumination conditions limits the speed and real-time applications. This paper introduces EventPS, a novel approach to real-time photometric stereo using an event camera. Capitalizing on the exceptional temporal resolution, dynamic range, and low bandwidth character-istics of event cameras, EventPS estimates surface nor-mal only from the radiance changes, significantly enhancing data efficiency. EventPS seamlessly integrates with both optimization-based and deep-learning-based photo-metric stereo techniques to offer a robust solution for non-Lambertian surfaces. Extensive experiments validate the effectiveness and efficiency of EventPS compared to frame-based counterparts. Our algorithm runs at over 30 fps in real-world scenarios, unleashing the potential of EventPS in time-sensitive and high-speed downstream applications.11Code available: https://codeberg.org/ybh1998/EventPS Bohan Yu, Jieji Ren, Jin Han 0001, Feishi Wang, Jinxiu Liang, Boxin Shi |
CVPR | 2 |
| 2024 | AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the WildabstractWhile humans can use parts of their arms other than the hands for manipulations like gathering and supporting, whether robots can effectively learn and perform the same type of operations remains relatively unexplored. As these manipulations require joint-level control to regulate the complete poses of the robots, we develop AirExo, a low-cost, adaptable, and portable dual-arm exoskeleton, for teleoperation and demonstration collection. As collecting teleoperated data is expensive and time-consuming, we further leverage AirExo to collect cheap in-the-wild demonstrations at scale. Under our in-the-wild learning framework, we show that with only 3 minutes of the teleoperated demonstrations, augmented by diverse and extensive in-the-wild data collected by AirExo, robots can learn a policy that is comparable to or even better than one learned from teleoperated demonstrations lasting over 20 minutes. Experiments demonstrate that our approach enables the model to learn a more general and robust policy across the various stages of the task, enhancing the success rates in task completion even with the presence of disturbances. Project website: airexo.github.io. Hongjie Fang, Haoshu Fang, Jieji Ren, Ruo Zhang, Cewu Lu |
ICRA | 4 |
| 2024 | Differentiable Fluid Physics Parameter Identification By Stirring and For StirringabstractFluid interactions are crucial in daily tasks, with properties like density and viscosity being key parameters. The property states can be used as control signals for robot operation. While density estimation is simple, assessing viscosity, especially for different fluid types, is complex. This study introduces a novel differentiable fitting framework, DiffStir, tailored to identify key physics parameters through stirring. Then, given the estimated physics parameters, we can generate commands to guide the robotic stirring. Comprehensive experiments were conducted to validate the efficacy of DiffStir, showcasing its precision in parameter estimation when benchmarked against reported values in the literature. More experiments and videos can be found in the supplementary materials and on the website: https://diffstir.robotflow.ai. Dongzhe Zheng, Yutong Li 0004, Jieji Ren, Cewu Lu |
IROS | 4 |
| 2024 | Visual-Language Collaborative Representation Network for Broad-Domain Few-Shot Image ClassificationabstractVisual-language models based on CLIP have shown remarkable abilities in general few-shot image classification. However, their performance drops in specialized fields such as healthcare or agriculture, because CLIP's pre-training does not cover all category data. Existing methods excessively depend on the multi-modal information representation and alignment capabilities acquired from CLIP pre-training, which hinders accurate generalization to unfamiliar domains. To address this issue, this paper introduces a novel visual-language collaborative representation network (MCRNet), aiming at acquiring a generalized capability for collaborative fusion and representation of multi-modal information. Specifically, MCRNet learns to generate relational matrices from an information fusion perspective to acquire aligned multi-modal features. This relationship generation strategy is category-agnostic, so it can be generalized to new domains. A class-adaptive fine-tuning inference technique is also introduced to help MCRNet efficiently learn alignment knowledge for new categories using limited data. Additionally, the paper establishes a new broad-domain few-shot image classification benchmark containing seven evaluation datasets from five domains. Comparative experiments demonstrate that MCRNet outperforms current state-of-the-art models, achieving an average improvement of 13.06% and 13.73% in the 1-shot and 5-shot settings, highlighting the superior performance and applicability of MCRNet across various domains. Jieji Ren, Haofen Wang, Tianxing Wu 0001, Weifeng Ge |
ACM Multimedia | 2 |
| 2024 | A Generalized Motion Control Framework of Dielectric Elastomer Actuators: Dynamic Modeling, Sliding-Mode Control and Experimental EvaluationabstractThe continuous electromechanical deformation of dielectric elastomer actuators (DEAs) suffers from rate-dependent viscoelasticity, mechanical vibration, and configuration dependency, making the generalized dynamic modeling and precise control elusive. In this work, we present a generalized motion control framework for DEAs capable of accommodating different configurations, materials and degrees of freedom (DOFs). First, a generalized, control-enabling dynamic model is developed for DEAs by taking both nonlinear electromechanical coupling, mechanical vibration and rate-dependent viscoelasticity into consideration. Further, a state observer is introduced to predict the unobservable viscoelasticity. Then, an enhanced exponential reaching law-based sliding-mode controller (EERLSMC) is proposed to minimize the viscoelasticity of DEAs. Its stability is also proved mathematically. The experimental results obtained for different DEAs (four configurations, two materials, and multi-DOFs) demonstrate that our dynamic model can precisely describe their complex dynamic responses and the EERLSMC can achieve precise tracking control; verifying the generality and versatility of our motion control framework. Shakiru Olajide Kassim, Jieji Ren, Vahid Vaziri, Sumeet S. Aphale, Guo-Ying Gu |
IEEE Trans. Robotics | 3 |
| 2023 | ReLeaPS : Reinforcement Learning-based Illumination Planning for Generalized Photometric StereoabstractIllumination planning in photometric stereo aims to find a balance between surface normal estimation accuracy and image capturing efficiency by selecting optimal light configurations. It depends on factors such as the unknown shape and general reflectance of the target object, global illumination, and the choice of photometric stereo backbones, which are too complex to be handled by existing methods based on handcrafted illumination planning rules. This paper proposes a learning-based illumination planning method that jointly considers these factors via integrating a neural network and a generalized image formation model. As it is impractical to supervise illumination planning due to the enormous search space for ground truth light configurations, we formulate illumination planning using reinforcement learning, which explores the light space in a photometric stereo-aware and reward-driven manner. Experiments on synthetic and real-world datasets demonstrate that photometric stereo under the 20-light configurations from our method is comparable to, or even surpasses that of using lights from all available directions. Jun Hoong Chan, Bohan Yu, Heng Guo 0003, Jieji Ren, Zongqing Lu 0002, Boxin Shi |
ICCV | 4 |
| 2023 | DiLiGenT-Π: Photometric Stereo for Planar Surfaces with Rich Details - Benchmark Dataset and BeyondabstractPhotometric stereo aims to recover detailed surface shapes from images captured under varying illuminations. However, existing real-world datasets primarily focus on evaluating photometric stereo for general non-Lambertian reflectances and feature bulgy shapes that have a certain height. As shape detail recovery is the key strength of photometric stereo over other 3D reconstruction techniques, and the near-planar surfaces widely exist in cultural relics and manufacturing workpieces, we present a new real-world dataset DiLiGenT-Π containing 30 near-planar scenes with rich surface details. This dataset enables us to evaluate recent photometric stereo methods specifically for their ability to estimate shape details under diverse materials and to identify open problems such as near-planar surface normal estimation from uncalibrated photometric stereo and surface detail recovery for translucent materials. To inspire future research, this dataset will open soruced at https://photometricstereo.github.io/diligentpi.html. Feishi Wang, Jieji Ren, Heng Guo 0003, Ming Jun Ren, Boxin Shi |
ICCV | 2 |
| 2023 | In-situ Mechanical Calibration for Vision-based Tactile SensorsabstractThis paper proposes a novel approach to conduct routine calibration for the changing mechanical parameters over time of a vision-based tactile sensor, without disassembling its overall structure, i.e., in-situ mechanical calibration. Calibration for mechanical parameters, Young's modulus and Poisson's ratio, of a tactile sensor's sensing elastomer, is crucial for its force perception capabilities. However, there are few methods that can retrieve values of these parameters both accurately and conveniently. To address this problem, we propose an in-situ approach to calibrate mechanical parameters other than the verbose traditional evaluation process. This method incorporates the deformation sensing capability of the sensor, the accurate force sensing capability of a force/torque sensor, and most importantly, the deformation-force relation-ship for an indentation with embedded mechanical parameters of the elastomers. We also present the indentation test setup and the complete pipeline to extract Young's modulus and Poisson's ratio from experimental results. We validate the method by comparing the indentation depths simulated through finite element analysis (FEA) using the cali-brated parameters with the indentation depths measured in real experiments. Furthermore, superior contact force distribution can be achieved with the accurate mechanical parameters. The proposed method provides the theoretical basis for accurate, lifelong routine calibration, whether weekly or even daily, which can enhance the applications of tactile sensors in real manipulation scenarios. Jieji Ren, Hexi Yu, Daolin Ma |
ICRA | 2 |
| 2022 | DiLiGenT102: A Photometric Stereo Benchmark Dataset with Controlled Shape and Material VariationabstractEvaluating photometric stereo using real-world dataset is important yet difficult. Existing datasets are insufficient due to their limited scale and random distributions in shape and material. This paper presents a new real-world photometric stereo dataset with “ground truth” normal maps, which is 10 times larger than the widely adopted one. More importantly, we propose to control the shape and material variations by fabricating objects from CAD models with carefully selected materials, covering typical aspects of reflectance properties that are distinctive for evaluating photometric stereo methods. By benchmarking recent photometric stereo methods using these 100 sets of images, with a special focus on recent learning based solutions, a 10x 10 shape-material error distribution matrix is visualized to depict a “portrait” for each evaluated method. From such comprehensive analysis, open problems in this field are discussed. To inspire future research, this dataset is available at https://photometricstereo.github.io. Jieji Ren, Feishi Wang, Ming Jun Ren, Boxin Shi |
CVPR | 1 |
| 2022 | Cooperative Light-Field Image Super-Resolution Based on Multi-Modality Embedding and Fusion With Frequency AttentionabstractLight field (LF) imaging is an advanced visual perception system, which can record the intensity and direction information of light rays and provide multi-viewpoint images from a single capture. However, there is a trade-off between spatial and angular resolutions due to the restricted sensor size, which limits the wide applications of LF cameras. To address this problem, we propose a cooperative network to super-resolve LF sub-aperture images based on the multi-modality fusion. Specifically, in order to fully explore the LF information, we adopt various modalities and extract corresponding features to emphasise diverse LF characteristics. Then, we design a multi-scale fusion module to effectively integrate global and local LF features and apply frequency-aware attention mechanism to adaptively reinforce fused features. Extensive experiments demonstrate the superiority of our method on both qualitative and quantitative evaluations, with competitive execution efficiency. Jieji Ren, Xiangchao Yan, Ming Jun Ren |
IEEE Signal Process. Lett. | 2 |
| 2020 | Multiscale Convolutional Fusion Network for Non-Lambertian Photometric StereoabstractOne of the key issues in photometric stereo is the extension of its application in real world objects which shows non-Lambertian reflectance. This letter proposes a multi-scale weighted convolutional fusion network with deep learning architecture to realize high-precision perception of non-Lambertian surfaces under arbitrary illumination conditions. A multi-scale convolutional fusion module is designed to strengthen the photometric physics and the utilization of the neighborhood features at the same time so as to overcome shadows and distinguish multiple materials. In order to further deal with the problem of arbitrary illumination conditions, a multi-resolution polar coordinate division method is proposed to integrate the input image information and fully utilizes the multi-scale convolution. Both syntheses and real-world experiments verifies the performance of the proposed method in recovery accuracy and computational efficiency. Jieji Ren, Xi Wang 0023, Zhenxiong Jian, Ming Jun Ren |
IEEE Signal Process. Lett. | 1 |