EDBT 2026 Demo / reviewers in the wild / expert
Xiaolong Shen
dblp:98/10204
· DBLP profile ↗
17ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Transient-Enhanced Buck Converter Based on Capacitor Current Adaptive On-Time (CCAOT) Control
Xiaolong Shen, Haolu Xie |
ISCAS | 1 |
| 2025 | A Causality-Driven Multi-Modal Alignment Method in Medical ClassificationabstractMedical images such as CT and PET play a significant role in cancer diagnosis, often accompanied by diagnostic reports that describe findings across multiple organs. These multimodal data sources offer clinicians detailed information to support comprehensive and accurate diagnoses. However, in the context of computer-aided diagnosis (CAD), the inherent noise and potential inconsistencies across different imaging modalities can introduce confounding information, thereby reducing diagnostic accuracy. To better extract disease-relevant features from full-body medical images and reports, we propose a causality-driven multi-modal alignment method. This approach extracts causal features from comprehensive medical images and structural reports, enabling machine learning models to focus on specific disease classification tasks. For both image and text modalities, we designed a causality-enhancement method to extract causal features from individual modalities respectively and align crossmodal features in fusion stage. Our approach provides a novel causal view on computer aided medical classification tasks. Xiaolong Shen, Yina Li, Xiaodong Yue 0002 |
BIBM | 1 |
| 2025 | Efficient Approximate Logic Synthesis with Dual-Phase Iterative FrameworkabstractApproximate computing is an emerging paradigm to improve the energy efficiency for error-tolerant applications. Many iterative approximate logic synthesis (ALS) methods were proposed to automatically design approximate circuits. However, as the sizes of circuits grow, the runtime of ALS grows rapidly. Thus, a crucial challenge is to ensure circuit quality while improving the efficiency of ALS. This work proposes a dual-phase iterative framework to accelerate the iterative ALS flows. In the first phase, a comprehensive circuit analysis is performed to gather the necessary information, including the error information. In the second phase, minimal incremental computation is employed based on the information from the first phase. The experimental results show that the proposed method achieves an acceleration by up to 21.8 × without loss of circuit quality compared to the state-of-the-art methods. Ruicheng Dai, Xuan Wang 0027, Wenhui Liang, Xiaolong Shen, Menghui Xu, Leibin Ni, Gezi Li, Weikang Qian |
DATE | 4 |
| 2025 | AccALS 2.0: Accelerating Approximate Logic Synthesis by Simultaneous Selection of Multiple Local Approximate ChangesabstractApproximate computing emerges as an energy-efficient computing paradigm designed for applications that can tolerate errors. Many iterative methods for approximate logic synthesis (ALS) have been developed to automatically synthesize approximate circuits. Nonetheless, most of them overlook the potential of applying multiple local approximate changes (LACs) simultaneously in one iteration, which can significantly reduce the overall computation time. In this article, we propose AccALS 2.0, a novel framework for further accelerating iterative ALS flows, which is based on simultaneous selection of multiple LACs in a single round. However, there are two challenges for selecting multiple LACs. The first is that the mutual influence of multiple LACs can affect the estimation of the circuit error. The second is that there may exist conflicts among multiple LACs. To address these issues, first, we propose an efficient measure for the mutual influence between two LACs. With its help, we transform the problems of solving the LAC conflicts and selecting multiple LACs into a unified maximum independent set problem for solving. The experimental results showed that AccALS 2.0 outperforms state-of-the-art ALS methods in runtime, while achieving similar or better-circuit quality. Xuan Wang 0027, Xiaomi Zhou, Ruicheng Dai, Xiaolong Shen, Menghui Xu, Leibin Ni, Weikang Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Controllable 3D Face Generation with Conditional Style Code DiffusionabstractGenerating photorealistic 3D faces from given conditions is a challenging task. Existing methods often rely on time-consuming one-by-one optimization approaches, which are not efficient for modeling the same distribution content, e.g., faces. Additionally, an ideal controllable 3D face generation model should consider both facial attributes and expressions. Thus we propose a novel approach called TEx-Face(TExt & Expression-to-Face) that addresses these challenges by dividing the task into three components, i.e., 3D GAN Inversion, Conditional Style Code Diffusion, and 3D Face Decoding. For 3D GAN inversion, we introduce two methods, which aim to enhance the representation of style codes and alleviate 3D inconsistencies. Furthermore, we design a style code denoiser to incorporate multiple conditions into the style code and propose a data augmentation strategy to address the issue of insufficient paired visual-language data. Extensive experiments conducted on FFHQ, CelebA-HQ, and CelebA-Dialog demonstrate the promising performance of our TEx-Face in achieving the efficient and controllable generation of photorealistic 3D faces. The code will be publicly available. Xiaolong Shen, Chang Zhou 0005, Zongxin Yang |
AAAI | 1 |
| 2024 | StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language RecognitionabstractThe goal of sign language recognition (SLR) is to help those who are hard of hearing or deaf overcome the communication barrier. Most existing approaches can be typically divided into two lines, i.e., Skeleton-based, and RGB-based methods, but both lines of methods have their limitations. Skeleton-based methods do not consider facial expressions, while RGB-based approaches usually ignore the fine-grained hand structure. To overcome both limitations, we propose a new framework called the Spatial-temporal Part-aware network (StepNet), based on RGB parts. As its name suggests, it is made up of two modules: Part-level Spatial Modeling and Part-level Temporal Modeling. Part-level Spatial Modeling, in particular, automatically captures the appearance-based properties, such as hands and faces, in the feature space without the use of any keypoint-level annotations. On the other hand, Part-level Temporal Modeling implicitly mines the long short-term context to capture the relevant attributes over time. Extensive experiments demonstrate that our StepNet, thanks to spatial-temporal modules, achieves competitive Top-1 Per-instance accuracy on three commonly used SLR benchmarks, i.e., 56.89% on WLASL, 77.2% on NMFs-CSL, and 77.1% on BOBSL. Additionally, the proposed method is compatible with the optical flow input and can produce superior performance if fused. For those who are hard of hearing, we hope that our work can act as a preliminary step. Xiaolong Shen, Zhedong Zheng, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Global-to-Local Modeling for Video-Based 3D Human Pose and Shape EstimationabstractVideo-based 3D human pose and shape estimations are evaluated by intra-frame accuracy and inter-frame smoothness. Although these two metrics are responsible for different ranges of temporal consistency, existing state-of-the-art methods treat them as a unified problem and use monotonous modeling structures (e.g., RNN or attention-based block) to design their networks. However, using a single kind of modeling structure is difficult to balance the learning of short-term and long-term temporal correlations, and may bias the network to one of them, leading to undesirable predictions like global location shift, temporal inconsistency, and insufficient local details. To solve these problems, we propose to structurally decouple the modeling of long-term and short-term correlations in an end-to-end framework, Global-to-Local Transformer (GLoT). First, a global transformer is introduced with a Masked Pose and Shape Estimation strategy for long-term modeling. The strategy stimulates the global transformer to learn more inter-frame correlations by randomly masking the features of several frames. Second, a local transformer is responsible for exploiting local details on the human mesh and interacting with the global transformer by leveraging cross-attention. Moreover, a Hierarchical Spatial Correlation Regressor is further introduced to refine intra-frame estimations by decoupled global-local representation and implicit kinematic constraints. Our GLoT surpasses previous state-of-the-art methods with the lowest model parameters on popular benchmarks, i.e., 3DPW, MPI-INF-3DHP, and Human3.6M. Codes are available at https://github.com/sxl142/GLoT. Xiaolong Shen, Zongxin Yang, Chang Zhou 0005, Yi Yang 0001 |
CVPR | 1 |
| 2023 | High-accuracy Low-power Reconfigurable Architectures for Decomposition-based Approximate Lookup TableabstractStoring pre-computed results of frequently-used functions into lookup table (LUT) is a popular way to improve energy efficiency, but its advantage diminishes as the number of input bits increases. A recent work shows that by decomposing the target function approximately, the total LUT entries can be dramatically reduced, leading to significant energy saving. However, its heuristic approximate decomposition algorithm leads to sub-optimal approximation quality. Also, its rigid hardware architecture only supports disjoint decomposition and may have unnecessary extra power consumption sometimes. To address these issues, we develop a novel approximate decomposition algorithm based on beam search and simulated annealing, which can reduce 11.1% approximation error. We also propose a non-disjoint approximate decomposition method and two reconfigurable architectures. The first has 10.4% less error using 19.2% less energy and the second has 23.0% less error with same energy consumption compared to the state-of-the-art design. Xingyue Qian, Chang Meng, Xiaolong Shen, Junfeng Zhao 0003, Leibin Ni, Weikang Qian |
DATE | 3 |
| 2022 | SEALS: sensitivity-driven efficient approximate logic synthesisabstractApproximate computing is an emerging computing paradigm to design energy-efficient systems. Many greedy approximate logic synthesis (ALS) methods have been proposed to automatically synthesize approximate circuits. They typically need to consider all local approximate changes (LACs) in each iteration of the ALS flow to select the best one, which is time-consuming. In this paper, we propose SEALS, a Sensitivity-driven Efficient ALS method to speed up a greedy ALS flow. SEALS centers around a newly proposed concept called sensitivity, which enables a fast and accurate error estimation method and an efficient method to filter out unpromising LACs. SEALS can handle any statistical error metric. The experimental results show that it outperforms a state-of-the-art ALS method in runtime by 12X to 15X without reducing circuit quality. Chang Meng, Xuan Wang 0027, Sijun Tao, Zhihang Wu, Leibin Ni, Xiaolong Shen, Junfeng Zhao 0003, Weikang Qian |
DAC | 8 |
| 2022 | VECBEE: A Versatile Efficiency-Accuracy Configurable Batch Error Estimation Method for Greedy Approximate Logic SynthesisabstractApproximate computing is an emerging strategy to improve the energy efficiency of many error-tolerant applications. To design an approximate circuit automatically, many approximate logic synthesis (ALS) methods have been proposed, among which many are greedy. To improve the synthesis quality of these greedy methods, one key is to calculate the errors of all candidate approximate transformations accurately. However, the traditional simulation-based method is time consuming. Instead, many existing methods just perform quick but inaccurate error estimation. In this work, to improve both the accuracy and runtime of error estimation, we propose VECBEE, a versatile efficiency–accuracy configurable batch error estimation method for greedy ALS. It is based on Monte Carlo simulation and an efficient technique to capture whether a signal change due to an introduced approximation will be propagated to each primary output. VECBEE is generally applicable to any statistical error measurement, such as error rate and average error magnitude, and any graph-based circuit representation. It allows a flexible tradeoff between the error estimation accuracy and the runtime, while even the fully accurate version is much faster than the traditional simulation-based method. We apply VECBEE to two representative greedy ALS methods and demonstrate its effectiveness in generating better approximate circuits. The code of VECBEE is made open source. Sanbao Su, Chang Meng, Fan Yang 0001, Xiaolong Shen, Leibin Ni, Zhihang Wu, Junfeng Zhao 0003, Weikang Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Multi/Many-Objective Optimization Via A New Preference Indicator
Lianbo Ma 0004, Mingli Shi, Rui Wang 0017, Shengminjie Chen, Junfei Zhao, Xiaolong Shen |
CEC | 6 |
| 2020 | Two-Stage Learning Brain Storm Optimizer
Junfeng Zhao 0003, Xiaolong Shen |
ICIC (1) | 5 |
| 2020 | A Modified Bacterial Foraging Optimizer with Adaptive Chemotactic Step in Dynamic Search Region
Yibo Yong, Junfeng Zhao 0003, Xiaolong Shen |
ICIC (1) | 4 |
| 2019 | Exploring frame segmentation networks for temporal action localization
Ke Yang 0004, Xiaolong Shen, Peng Qiao, Shijie Li 0002, Dongsheng Li 0001, Yong Dou |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Distributed sparse bundle adjustment algorithm based on three-dimensional point partition and asynchronous communicationabstractSparse bundle adjustment (SBA) is a key but time- and memory-consuming step in three-dimensional (3D) reconstruction. In this paper, we propose a 3D point-based distributed SBA algorithm (DSBA) to improve the speed and scalability of SBA. The algorithm uses an asynchronously distributed sparse bundle adjustment (A-DSBA) to overlap data communication with equation computation. Compared with the synchronous DSBA mechanism (SDSBA), A-DSBA reduces the running time by 46%. The experimental results on several 3D reconstruction datasets reveal that our distributed algorithm running on eight nodes is up to five times faster than that of the stand-alone parallel SBA. Furthermore, the speedup of the proposed algorithm (running on eight nodes with 48 cores) is up to 41 times that of the serial SBA (running on a single node). Xiaolong Shen, Yong Dou, Steven Mills, David M. Eyers, Huan Feng, Zhiyi Huang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2016 | Coarse-Grained Architecture for Fingerprint MatchingabstractFingerprint matching is a key procedure in fingerprint identification applications. The minutiae-based fingerprint matching algorithm is one of the most typical algorithms achieving a reasonably correct recognition rate. This study proposes a coarse-grained parallel architecture called fingerprint matching core (FMC) to accelerate fingerprint matching. The proposed architecture has a two-level parallel structure (i.e., parallel among groups (PAG) and parallel in group (PIG)). A multirequest controller is added to the PAG structure to obtain a concurrent operation of the multiple processing element group (PEG). The DDR3 controller is used in the PIG structure to read eight minutiae from eight different fingerprints and realize the simultaneous computation of the eight PEs. The whole system is implemented on a Xilinx FPGA board with a Virtex VII XC7VX485T chip. The 16-PEG FMC achieves a throughput of about 9.63 million fingerprint pairs per second, which is larger than that achieved on a Tesla K20c platform. The software execution times are also measured on the 2.93GHz Intel Xeon 5670, 2.3GHz AMD Opteron(tm) Processor 6376, and Tesla K20c platforms. The Intel Xeon 5670 has two processors with 12 cores, and the AMD Opteron(tm) Processor 6376 has two processors with 16 cores. Moreover, the throughput is about 31 times that achieved on a 2.93GHz Intel Xeon 5670 single core. Jinwei Xu, Jingfei Jiang, Yong Dou, Xiaolong Shen |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2014 | Indoor positioning fusion algorithm for smartphonesabstractIn this paper, a particle filter based fusion algorithm running on smartphones is introduced. It adopts signals, camera and IMU to achieve an accurate indoor positioning. The fusion algorithm achieves a localization accuracy of 2 meter during 80% test time and the maximum error is 4.5 meter. Xiaolong Shen |
IPIN | 3 |