EDBT 2026 Demo / reviewers in the wild / expert
Changlin Chen
dblp:123/8925
· DBLP profile ↗
17ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoupling Control of a Multi-Segment Hybrid-Actuated Soft Origami Continuum Robot Through Variable StiffnessabstractSoft origami continuum robots have attracted considerable attention because of their flexibility and adaptability, but they face challenges in control accuracy. This study presents a novel design and control scheme for large-deformation hybrid-actuated soft origami continuum robots to improve their motion performance in real tasks. An origami-based pneumatic chamber is used as the robot backbone to achieve a high extension ratio, and the hybrid variable stiffness principle combining antagonistic actuation and layer jamming further enhances the robot's overall structural stiffness. The robot performs precise movements by controlling tendons distributed in an external origami. An iterative training strategy based on long short-term memory is employed to model the inverse kinematics of the robot in consideration of the inherent hysteresis of origami robots. Step size features are introduced to improve model accuracy with limited data. The capability for variable stiffness enables the migration of the training model on the basis of a single segment, the accuracy of which is close to the submillimeter level. Decoupling control of each segment for a rear-driven multi-segment origami continuum robot is also achieved. Experiment results reveal that the trajectory tracking errors for single-segment and multi-segment of the robot are 1.58, and 2.81 mm, respectively, with relative errors of 0.71% and 0.63% over the robot length, demonstrating the good performance of the proposed design and control method. Orientation data can also be added to the dataset to achieve orientation control of the robot, with an angle error of 0.75°. The robot shows its stability and safety in a series of real tasks, including writing, LED tracking, board cleaning, and pick and place. This study demonstrates a potential solution for continuum robots performing precise tasks without any sensory feedback in confined spaces and specialized environments where electronic sensors fail. Tianheng Li, Changlin Chen, Qiqiang Hu, Erbao Dong, Shiwu Zhang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Chain-of-Detection: Enhancing Cross-Granularity Robotic Perception for Object ManipulationabstractIn robotic perception, cross-granularity object detection is essential for identifying and localizing targets at varying levels of detail. Traditional detection methods often struggle to bridge the gap between coarse object detection and fine-grained component localization, limiting their ability to associate parts, such as a cup and its handle. Vision-language models (VLMs), while effective in spatial reasoning, face challenges in fine-grained detection due to the scarcity of annotated datasets. To address these issues, we first propose the chain-of-detection (CoD) framework, which focuses on guiding detection in a step-by-step manner from coarse recognition to fine-grained localization. During this process, we observe that existing detectors still lack sufficient capability in recognizing fine-grained components. To overcome this limitation, we further combine the CoD framework with Monte Carlo tree search (MCTS) to automatically generate fine-grained datasets, eliminating the need for manual labeling and significantly improving detector performance. Experiments show that our approach achieves an average improvement of 17.31% in robotic manipulation success rates for common objects, 51.39% for larger object operations, and about 50% in simulated environments. These results demonstrate the effectiveness of CoD in advancing cross-granularity detection and enhancing precise robotic manipulation. The implementation is publicly available at https://github.com/tinnel123666888/CoD and the CoD dataset is released at https://huggingface.co/datasets/tinnel123/CoD_dataset. Tianrun Xu, Haichuan Gao, Changlin Chen, Shiyuan Xu, Shangqi Guo, Feng Chen 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-Based NN AcceleratorabstractCompute-in-Memory (CIM) and weight sparsity are two effective techniques to reduce data movement during Neural Network (NN) inference. However, they can hardly be employed in the same accelerator simultaneously because CIM requires structural compute patterns which are disrupted in sparse NNs. In this paper, we partially solve this issue by proposing a bit level weight reordering strategy which can realize compact mapping of sparse NN weight matrices onto Resistive Random Access Memory (RRAM) based NN Accelerators (RRAM-Acc). In specific, when weights are mapped to RRAM crossbars in a binary complement manner, we can observe that, which can also be mathematically proven, bit-level sparsity and similarity commonly exist in the crossbars. The bit reordering method treats bit sparsity as a special case of bit similarity, reserve only one column in a pair of columns that have identical bit values, and then map the compressed weight matrices into Operation Units (OU). The performance of our design is evaluated with typical NNs. Simulation results show a 61.24 % average performance improvement and$1.51 \times-2.52 \times$energy savings under different sparsity ratios, with only slight overhead compared to the state-of-the-art design. Weiping Yang, Shilin Zhou 0001, Yujiao Nie, Qimin Zhou, Changlin Chen |
ICPADS | 7 |
| 2025 | ICBSS: An Improved Algorithm for Multi-Agent Combinatorial Path FindingabstractThe Multi-Agent Combinatorial Path Finding (MCPF) problem is a generalized version of the Multi-Agent Path Finding (MAPF) problem, in which each agent must collectively visit multiple intermediate target locations on the way to its final destination. The state-of-the-art approach for addressing MCPF, known as Conflict-Based Steiner Search (CBSS) [1], leverages K-best joint sequences to create multiple search trees, and employs a CBS-like search to resolve collisions for each tree. Despite its optimality guarantee, CBSS is computationally burdensome due to the duplicated collision resolutions across multiple trees and the computation of the K best joint sequences. To address these challenges, we propose a novel algorithm called Improved Conflict-Based Steiner Search (ICBSS), aiming at expediting CBSS by replacing the multi trees with a single constraint tree (CT), which can be implemented by interleaving the time-dependent traveling salesman algorithm to compute the optimal joint path for agents under the newly generated constraints in each CT vertex. Additionally, we introduce a sub-optimal variant of ICBSS, which improves computational efficiency at the expense of solution optimality. Empirical results show that ICBSS outperforms state-of-the-art MCPF algorithms on a variety of MAPF instances. Zheng Chen 0004, Changlin Chen, Yiran Ni |
ICRA | 2 |
| 2025 | Heuristically Guided Compilation for Task Assignment and Path FindingabstractWe investigate the Combined Target-Assignment and Path-Finding (TAPF) problem that computes both task assignments and collision-free paths for multiple agents, that is, each agent is required to select a target from an underlying set, reaching which leads to a payoff. There is a cost closely related to the time required for each agent to reach the goal. The objective is to maximize the minimum gain generated by the agents. We proposed a Compilation-Based Approach with Heuristics (TA-CBWH) to approximate the optimal solution, behind which are two critical ideas: (i) for a specific task assignment, we formulate an integer linear programming (ILP) and create the iteration combined with large neighborhood search (LNS) to quickly improve the solution quality to near-optimal; (ii) regarding distinct task assignments, a switching mechanism is developed to determine the most promising iteration while progressively eliminating unnecessary task assignments. Comparative experiments demonstrate that TA-CBWH outperforms a wide range of existing approaches across various maps and different numbers of agents. Changlin Chen, Yiran Ni |
ICRA | 2 |
| 2025 | OPASCA: Outer Product-Based Accelerator With Unified Architecture for Sparse Convolution and AttentionabstractVision transformer (ViT)-based models have achieved state-of-the-art accuracy in many computer vision tasks, but their attention mechanism is more computation and communication intensive than convolutional neural networks (CNNs). To adapt ViT-based models for resource-constrained edge computing platforms, techniques, such as network sparsity and convolution-attention combination, have been proposed to reduce processing costs without compromising accuracy. Therefore, specific hardware designs able to handle sparsity in both convolution and attention operations are required to accelerate the processing speed. In view of this, this work proposes OPASCA, an accelerator featuring a unified hardware architecture that supports irregular activation and weight sparsity in both operations. Specifically, this target is achieved with the following contributions: 1) we employ outer product dataflow to efficiently handle sparse weights and input neurons, supporting both convolution and attention computing with minimal hardware overhead; 2) we design a hierarchical butterfly network to route the output neurons to the accumulation buffers, minimizing the conflicts among outer product results and reducing the hardware overhead of accumulation banks; and 3) we propose a novel encoding scheme that achieves more compact sparse inputs and enhances multiplier utilization. Evaluations on VGG-16, ViT, BoTNet, and Conformer models show that OPASCA outperforms state-of-the-art accelerators by$1.60 \times -2.08 \times $on sparse convolution tasks and by$1.18 \times -1.70 \times $on sparse attention tasks in term of performance. It also reduces DRAM access to 19%–86% and energy consumption to 36%–92% of those of the counterpart designs. Qimin Zhou, Tuo Ma, Changlin Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Estate: Expert-Guided State Text Enhancement for Zero-Shot Industrial Anomaly DetectionabstractThe Expert-Guided State Text Enhancement Anomaly Detection (ESTATE) framework addresses the challenges in industrial anomaly detection arising from diverse product categories and limited defective samples. This framework, integrating expert insights through comparative state prompts, leverages two innovative text-guided networks, CLS-Refiner and SEG-Refiner, enhancing model training. These networks, connected to residual textual features of standard vision-language pre-trained models, focus on amplifying adjectives’ significance in text for improved image block and pixel-level alignment. ESTATE’s effectiveness is demonstrated through evaluations on MVTecAD and VisA datasets, achieving AUROC scores of 89.6%/89.6% for classification and 95.1%/85.0% for segmentation tasks, alongside setting new benchmarks in F1Max and PRO metrics. The AUC-cls on MVTecAD and VisA demonstrated an enhancement of 5.06% and 8.97%, respectively, compared to the APRIL-GAN approach. Bingke Zhu, Hao Li 0115, Changlin Chen, Liujie Hua, Jinqiao Wang |
ICIP | 3 |
| 2024 | An Integration and Time-Sampling based Readout Circuit with Current Compensation for Parallel MAC operations in RRAM ArraysabstractIn Resistive Random Access Memory (RRAM) based compute-in-memory designs, the column current readout circuits still consume too much area and power overhead, even if plenty of methods have been proposed to optimize the circuits. To alleviate this problem, this paper presents a novel current readout circuit to sense multiply-and-accumulate (MAC) result of RRAM array. Specifically, the circuit first integrate the stabilized and proportionally mirrored MAC current on a small capacitor until it fires, then sample the integration time with a set of reference signals with different carefully designed delays, and finally code the sampled result into a digital value. The proposed readout circuit has fine stability due to simple and determined relationship among the inputs, the RRAM cells’ states, and the MAC current. Meanwhile, the proposed design can achieve accurate MAC result readout at low resistance switching ratios, for it employs a current compensation circuit to remove background current caused by high resistance state RRAM cells. Our design is implemented using 28nm CMOS technology with a read latency of 2.8ns and an area occupation of 267μm2/channel, which is 12.5% and 89% less than state of the art design. Its power consumption, 0.092mW/channel, is also less than most counterpart designs. Weiping Yang, Shilin Zhou 0001, Qimin Zhou, Qingjiang Li, Changlin Chen |
ISCAS | 8 |
| 2024 | Regularized joint self-training: A cross-domain generalization method for image classification
Changlin Chen, Minghao Liu 0017, Zhaomin Rong, Shuangbao Shu |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | A storage-efficient SNN-CNN hybrid network with RRAM-implemented weights for traffic signs recognition
Yufei Zhang 0003, Lixing Huang, Changlin Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Conversion of infrared ocean target images to visible images driven by energy information
Changlin Chen, Xuewei Chao |
Multim. Syst. | 1 |
| 2023 | Few-shot ship classification based on metric learning
Changlin Chen, Shukun Ma |
Multim. Syst. | 2 |
| 2018 | FPGA-Based Multi-core Reconfigurable System for SAR ImagingabstractWith the development of the very large scale integration circuit (VLSI), multi-core processors have been developing fast and provide a novel approach to meet the requirements of modern complex computing. Real-time SAR imaging is always a hard problem to deal with since the huge amount of data and complex processing. Therefore, we design a schedulable and scalable multi-core parallel architecture based on FPGA and map the fundamental Chirp Scaling algorithm to the system. The design of the master control core makes the system highly extensible and the dedicated computing cores are designed to accelerate the process. In addition, the system can meet the real-time requirements, and it performs well in the resource utilization and the FPGA chip takes up a small space as well. Wei Di, Changlin Chen, Yongxiang Liu |
IGARSS | 2 |
| 2017 | Towards Maximum Utilization of Remained Bandwidth in Defected NoC LinksabstractTo maximize the utilization of the available networks-on-chip (NoCs) link bandwidth, partially faulty links with low fault level should be utilized while heavily defected (HD) links should be deactivated and dealt with by means of a fault tolerant routing algorithm. To reach this target, we make the following contributions in this paper: 1) we propose a flit serialization (FS) method to efficiently utilize partially faulty links. The FS approach divides the links into a number of equal width sections, and serializes sections of adjacent flits to transmit them on all fault-free link sections to mitigate the unbalance between the flit size and the actual link bandwidth; 2) we propose the link augmentation with one redundant section as a low cost mechanism to mitigate the FS drawback that a link's available bandwidth is reduced even if it contains only one faulty wire; and 3) we deactivate HD links when their fault level exceed a certain threshold to diminish congestion caused by HD links. The optimal threshold is derived by comparing the zero load packet transmission latency on the HD links and that on the shortest alternative path. Our proposal is evaluated with synthetic traffic and PARSEC benchmarks. Experimental results indicate that the FS method can achieve lower area*power/saturation_throughput value than all state of the art link fault tolerant strategies. With a redundant section in each link, the NoC saturation throughput can be largely improved than just utilizing FS, e.g., 18% when 10% of the NoC wires are broken. Simulation results we obtained at various wire broken rate configurations indicate that we achieve the highest saturation throughput if 4- or 8-section links with a flit transmission latency longer than four cycles are deactivated. Changlin Chen, Yaowen Fu, Sorin Cotofana |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Enabling vertical wormhole switching in 3D NoC-bus hybrid systems
Changlin Chen, Marius Enachescu, Sorin Cotofana |
DATE | 1 |
| 2013 | An Effective Routing Algorithm to Avoid Unnecessary Link Abandon in 2D Mesh NoCsabstractIn NoCs where each interconnection between neighboring routers is composed of a pair of unidirectional links, a broken link usually leads to the abandon of the entire interconnection, even if the other one is still functional. In this paper, we propose a fault tolerant Routing Algorithm (RA) which can efficiently utilize these fault free links when their pair broken links have available misrouting-contour sides. Constraints on the usage of Virtual Channels are adaptively applied according to the fault distribution, to avoid deadlock and unnecessary resource reservation. When compared with solid fault region tolerant RAs, which always abandon the entire interconnection, the proposed algorithm has twice higher saturation point under synthetic uniform traffics, and can on average diminish the execution time overhead for the evaluated applications, sample and sat ell, by 62.6% and 76.6%, respectively. Our experiments indicate that the embedding of the proposed algorithm into a baseline router increases the area cost and power consumption by 7.43% and 4.43%, respectively, which is not that significant given that the platform area is usually dominated by the computing cores area. Changlin Chen, Sorin Cotofana |
DSD | 1 |
| 2012 | A Novel Flit Serialization Strategy to Utilize Partially Faulty Links in Networks-on-ChipabstractAggressive MOS transistor size scaling substantially increase the probability of faults in NoC links due to manufacturing defects, process variations, and chip wire-out effects. Strategies have been proposed to tolerate faulty wires by replacing them with spare ones or by partially using the defective links. However, these strategies either suffer from high area and power overheads, or significantly increase the average network latency. In this paper, we propose a novel flit serialization method, which divides the links and flits into several sections, and serializes flit sections of adjacent flits to transmit them on all available fault-free link sections to avoid the complete waste of defective links bandwidth. Experimental results indicate that our method reduces the latency overhead significantly and enables graceful performance degradation, when compared with related partially faulty link usage proposals, and saves area and power overheads by up to 29% and 43.1%, respectively, when compared with spare wire replacement methods. Changlin Chen, Ye Lu 0003, Sorin Cotofana |
NOCS | 1 |