EDBT 2026 Demo / reviewers in the wild / expert
Yong Ding 0003
dblp:45/4509-3
· DBLP profile ↗
25ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-5226-7511ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | O3DAA: A method for offline 3D point cloud automatic annotation
Yong Ding 0003 |
Pattern Recognit. Lett. | 1 |
| 2024 | Improving Radial Imbalances with Hybrid Voxelization and RadialMix for LiDAR 3D Semantic SegmentationabstractHuge progress has been made in LiDAR 3D semantic segmentation, but there are two under-explored imbalances on the radial axis: points are unevenly concentrated on the near side, and the distribution of foreground object instances is skewed to the near side. This leads the training of the model to favor semantics at the near side with the majority of points and object instances. Both the cylindrical and the spherical voxelizations aim to address the problem of imbalanced point distribution by increasing the volume of voxels along the radial distance to include fewer near-side points in a smaller voxel and more far-side points in a bigger voxel. However, this causes a problem of the receptive field enlarging along the radial distance, which is not desirable in LiDAR 3D segmentation. This can be addressed in cubic voxelization which has a fixed volume of voxels. Thus, we propose a new LiDAR 3D semantic segmentation network (Hi-VoxelNet) with Hybrid Voxelization that leverages the advantages of cubic, cylindrical, and spherical voxelizations for hybrid voxel feature learning. To address the radial imbalance of object instances, we propose a novel data augmentation technique termed as RadialMix that uses radial sample duplication to increase the number of distant foreground object instances and mixes the radial duplication with another point cloud for enriching the training samples. With the joint improvements of the radial imbalances, our method archives state-of-the-art performance on nuScenes and SemanticKITTI datasets, and it shows significant improvements along the radial axis. Our code is publicly available at https://github.com/jialeli1/lidarseg3d. Hang Dai, Yu Wang 0177, Guangzhi Cao, Chun Luo, Yong Ding 0003 |
ICRA | 6 |
| 2024 | A 60-nA IQ 96.5% Peak Efficiency Buck Converter with Wide Load Range for Internet of ThingsabstractThis paper presents a 96.5% peak efficiency adaptive ON-time controlled buck converter with 0.001-500 mA load range for internet of things. A novel positive feedback controlled comparator is presented, with its bias current dynamically increased to accelerate the transient responses while maintaining low power consumption. Besides, the power gating scheme and a sub-nA voltage reference is implemented to achieve ultra-low overall quiescent current, leading to an outstanding conversion efficiency. Designed and simulated in a 0.18μm CMOS process, the prototype shows an undershoot voltage of 43 mV and recovery time of 2.8 µs during a load transient from 10 μA to 500 mA, with a total quiescent current of 60 nA. The proposed design shows efficiencies of 95.4%, 96.5% and 91.2% when the load current is 10 μA, 100 μA, and 500 mA, respectively. Ruijie Xi, Chenkang Xue, Jiping Li, Tianting Zhao, Mengjiao Li, Yong Ding 0003, Wuhua Li, Wanyuan Qu |
ISCAS | 6 |
| 2024 | LSOIT: Lexicon and Syntax Enhanced Opinion Induction Tree for Aspect-based Sentiment Analysis
Haiyan Wu, Di Zhou 0009, Chaoqun Sun, Zhiqiang Zhang 0010, Yong Ding 0003 |
Expert Syst. Appl. | 5 |
| 2024 | A Compact Dickson Hybrid Boost Converter With 5-mV Input 90.5% Peak Efficiency and On-Chip Cold-Start for Thermoelectric Energy HarvestingabstractThis article presents a novel Dickson hybrid boost converter for thermoelectric energy harvesting that utilizes a fixed ratio switched-capacitor (SC) converter combined with a traditional switched-inductor design. This unique approach makes high conversion ratio (CR) more attainable. Besides, the converter features an on-chip LC oscillator for cold-starting the harvester without requiring any external components. To improve the efficiency of the harvester, a lookup table-based maximum power point tracking (MPPT) and a zero current detection (ZCD) circuit are also implemented. The prototype was designed and fabricated using the 55 nm low power (LP) CMOS process, with a compact 5.6-$\mu $H inductor. Measurement results verify that the harvester can cold start at an open-circuit voltage ofV$_{\mathbf{TG}}$$=$110mV, achieve a peak efficiency of 90.5% atV$_{\mathbf{TG}}$$=$140mV, and maintain operation at a 5-mV input voltage with a 200-nA load. Furthermore, the converter maintains end-to-end efficiencies above 80% for a wide range ofV$_{\mathbf{TG}}$$=$40$\sim$200mV. With a wide input range of 5$\sim$100mV and load range of 200nA$\sim$4.52mA, the proposed converter represents a significant improvement in performance over previous state-of-the-art designs, with a decent power efficiency, compact form factor and wide output power range. Chenkang Xue, Ruijie Xi, Huipeng Xu, Lijie Shao, Xu Yang 0035, Yong Ding 0003, Wuhua Li, Wanyuan Qu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | CTVSR: Collaborative Spatial-Temporal Transformer for Video Super-ResolutionabstractVideo super-resolution (VSR) is important in video processing for reconstructing high-definition image sequences from corresponding continuous and highly-related video frames. However, existing VSR methods have limitations in fusing spatial-temporal information. Some methods only fuse spatial-temporal information on a limited range of total input sequences, while others adopt a recurrent strategy that gradually attenuates the spatial information. While recent advances in VSR utilize Transformer-based methods to improve the quality of the upscaled videos, these methods require significant computational resources to model the long-range dependencies, which dramatically increases the model complexity. To address these issues, we propose a Collaborative Transformer for Video Super-Resolution (CTVSR). The proposed method integrates the strengths of Transformer-based and recurrent-based models by concurrently assimilating the spatial information derived from multi-scale receptive fields and the temporal information acquired from temporal trajectories. In particular, we propose a Spatial Enhanced Network (SEN) with two key components: Token Dropout Attention (TDA) and Deformable Multi-head Cross Attention (DMCA). TDA focuses on the key regions to extract more informative features, and DMCA employs deformable cross attention to gather information from adjacent frames. Moreover, we introduce a Temporal-trajectory Enhanced Network (TEN) that computes the similarity of a given token with temporally-related tokens in the temporal trajectory, which is different from previous methods that evaluate all tokens within the temporal dimension. With comprehensive quantitative and qualitative experiments on four widely-used VSR benchmarks, the proposed CTVSR achieves competitive performance with relatively low computational consumption and high forward speed. Chenyan Lu, Zhengxue Liu, Hang Dai, Yong Ding 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingabstractLiDAR and camera are two modalities available for 3D semantic segmentation in autonomous driving. The popular LiDAR-only methods severely suffer from inferior segmentation on small and distant objects due to insufficient laser points, while the robust multi-modal solution is under-explored, where we investigate three crucial inherent difficulties: modality heterogeneity, limited sensor field of view intersection, and multi-modal data augmentation. We propose a multi-modal 3D semantic segmentation model (MSeg3D) with joint intra-modal feature extraction and inter-modal feature fusion to mitigate the modality heterogeneity. The multi-modal fusion in MSeg3D consists of geometry-based feature fusion GF-Phase, cross-modal feature completion, and semantic-based feature fusion SF-Phase on all visible points. The multi-modal data augmentation is reinvigorated by applying asymmetric transformations on LiDAR point cloud and multi-camera images individually, which benefits the model training with diversified augmentation transformations. MSeg3D achieves state-of-the-art results on nuScenes, Waymo, and SemanticKITTI datasets. Under the malfunctioning multi-camera input and the multi-frame point clouds input, MSeg3D still shows robustness and improves the LiDAR-only baseline. Our code is publicly available at https://github.com/jialeli1/lidarseg3d. Hang Dai, Yong Ding 0003 |
CVPR | 4 |
| 2023 | A 36-55 V Input 0.6-2.5 V Output Bypass-Assist Series-Capacitor Power Converter With 93.1% Peak Efficiency and 1.5 mA-5 A Load RangeabstractThis article presents a wide voltage and current range bypass-assist series-capacitor power converter for data centers and automotive applications. The overall power converter consists of a 2:1 charge pump (CP) front stage and a 6-level bypass-assist series-capacitor rear stage, by which the proposed design resolves the troublesome duty cycle limit that a typical series-capacitor converter usually faces. Therefore, the output regulation range is significantly extended and only one inductor is required. Besides, since the converter operates like a simple Buck converter with an equivalent input of$V_{\mathrm {IN}}$/12, an easy pulse-width-modulation (PWM) and pulse-frequency-modulation (PFM) compatible design is implemented and the efficiency greatly enhanced over a wide load range. The prototype was fabricated using the 180 nm bipolar-CMOS-DMOS (BCD) process and demonstrates a broad operating voltage range of 36~55-V input and 0.6~2.5-V output with 1.5mA~5A load range. For a typical 48-V input, the measured peak efficiencies of 2.5, 1.8, 1 and 0.6-V output are 93.1%, 88.1%, 86.6% and 82.7%, respectively. Additionally, when the load decreases to 30-mA, the proposed converter still achieves efficiencies of 88%, 82%, 80% and 70%, respectively. Chenkang Xue, Bingruo Gong, Huipeng Xu, Yanzheng Yu, Yong Ding 0003, Wuhua Li, Wanyuan Qu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Pseudo-Stereo for Monocular 3D Object Detection in Autonomous DrivingabstractPseudo-LiDAR 3D detectors have made remarkable progress in monocular 3D detection by enhancing the capability of perceiving depth with depth estimation networks, and using LiDAR-based 3D detection architectures. The Advanced stereo 3D detectors can also accurately localize 3D objects. The gap in image-to-image generation for stereo views is much smaller than that in image-to-LiDAR generation. Motivated by this, we propose a Pseudo-Stereo 3D detection framework with three novel virtual view generation methods, including image-level generation, feature-level generation, and feature-clone, for detecting 3D objects from a single image. Our analysis of depth-aware learning shows that the depth loss is effective in only feature-level virtual view generation and the estimated depth map is effective in both image-level and feature-level in our framework. We propose a disparity-wise dynamic convolution with dynamic kernels sampled from the disparity feature map to filter the features adaptively from a single image for generating virtual image features, which eases the feature degradation caused by the depth estimation errors. Till submission (November 18, 2021), our Pseudo-Stereo 3D detection framework ranks 1 st on car, pedestrian, and cyclist among the monocular 3D detectors with publications on the KITTI-3D benchmark. The code is released at https://github.com/revisitq/Pseudo-Stereo-3D. Yi-Nan Chen, Hang Dai, Yong Ding 0003 |
CVPR | 3 |
| 2022 | Cross-Modality Knowledge Distillation Network for Monocular 3D Object Detection
Hang Dai, Yong Ding 0003 |
ECCV (10) | 3 |
| 2022 | Self-Distillation for Robust LiDAR Semantic Segmentation in Autonomous Driving
Hang Dai, Yong Ding 0003 |
ECCV (28) | 3 |
| 2022 | A Dickson Hybrid Boost Converter With On-Chip Cold-Start for Thermoelectric Energy HarvestingabstractThis paper presents a Dickson hybrid boost converter with an on-chip cold starter for thermoelectric energy harvesting. The proposed harvester achieves both high power efficiency and low inductor volume. An on-chip LC oscillator cold starts the harvester without external components. Besides, a lookup table based maximum power point tracking (MPPT) improves the end-to-end efficiency in a wide input range. Designed and simulated in a 0.13 $\mu$m CMOS process, the harvester achieves a maximum end-to-end efficiency of 93.9% with a 5.6 $\mu$H inductor. The circuit cold starts at an input of 70 mV and the minimum operation voltage is as low as 6 mV with a 100 nW load. Chenkang Xue, Lijie Shao, Linhu Zhao, Xu Yang 0035, Yuqiu Lin, Yong Ding 0003, Wuhua Li, Wanyuan Qu |
ISCAS | 7 |
| 2021 | M3DSSD: Monocular 3D Single Stage Object DetectorabstractIn this paper, we propose a Monocular 3D Single Stage object Detector (M3DSSD) with feature alignment and asymmetric non-local attention. Current anchor-based monocular 3D object detection methods suffer from feature mismatching. To overcome this, we propose a two-step feature alignment approach. In the first step, the shape alignment is performed to enable the receptive field of the feature map to focus on the pre-defined anchors with high confidence scores. In the second step, the center alignment is used to align the features at 2D/3D centers. Further, it is often difficult to learn global information and capture long-range relationships, which are important for the depth prediction of objects. Therefore, we propose a novel asymmetric non-local attention block with multi-scale sampling to extract depth-wise features. The proposed M3DSSD achieves significantly better performance than the monocular 3D object detection methods on the KITTI dataset, in both 3D object detection and bird’s eye view tasks. The code is released at https://github.com/mumianyuxin/M3DSSD. Shujie Luo, Hang Dai, Ling Shao 0001, Yong Ding 0003 |
CVPR | 4 |
| 2021 | Anchor-free 3D Single Stage Detector with Mask-Guided Attention for Point CloudabstractMost of the existing single-stage and two-stage 3D object detectors are anchor-based methods, while the efficient but challenging anchor-free single-stage 3D object detection is not well investigated. Recent studies on 2D object detection show that the anchor-free methods also are of great potential. However, the unordered and sparse properties of point clouds prevent us from directly leveraging the advanced 2D methods on 3D point clouds. We overcome this by converting the voxel-based sparse 3D feature volumes into the sparse 2D feature maps. We propose an attentive module to fit the sparse feature maps to dense mostly on the object regions through the deformable convolution tower and the supervised mask-guided attention. By directly regressing the 3D bounding box from the enhanced and dense feature maps, we construct a novel single-stage 3D detector for point clouds in an anchor-free manner. We propose an IoU-based detection confidence re-calibration scheme to improve the correlation between the detection confidence score and the accuracy of the bounding box regression. Our code is publicly available at https://github.com/jialeli1/MGAF-3DSSD. Hang Dai, Ling Shao 0001, Yong Ding 0003 |
ACM Multimedia | 4 |
| 2021 | From Voxel to Point: IoU-guided 3D Object Detection for Point Cloud with Voxel-to-Point DecoderabstractIn this paper, we present an Intersection-over-Union (IoU) guided two-stage 3D object detector with a voxel-to-point decoder. To preserve the necessary information from all raw points and maintain the high box recall in voxel based Region Proposal Network (RPN), we propose a residual voxel-to-point decoder to extract the point features in addition to the map-view features from the voxel based RPN. We use a 3D Region of Interest (RoI) alignment to crop and align the features with the proposal boxes for accurately perceiving the object position. The RoI-Aligned features are finally aggregated with the corner geometry embeddings that can provide the potentially missing corner information in the box refinement stage. We propose a simple and efficient method to align the estimated IoUs to the refined proposal boxes as a more relevant localization confidence. The comprehensive experiments on KITTI and Waymo Open Dataset demonstrate that our method achieves significant improvements with novel architectures against the existing methods. The code is available on Github URLhttps://github.com/jialeli1/From-Voxel-to-Point . Hang Dai, Ling Shao 0001, Yong Ding 0003 |
ACM Multimedia | 4 |
| 2020 | Learning Local Quality-Aware Structures of Salient Regions for Stereoscopic Images via Deep Neural NetworksabstractThe perceptual quality of stereoscopic images plays an essential role in the human perception of visual information. However, most available stereoscopic image quality assessment (SIQA) methods evaluate 3D visual experience using hand-crafted features or shallow architectures, which cannot model the visual properties of stereo images well. In this paper, we use convolutional neural networks (CNNs) to learn deeper local quality-aware structures for stereo images. With different inputs, two CNN models are designed for no-reference SIQA tasks. The one-column CNN model directly accepts a cyclopean view as the input, and the three-column CNN model jointly considers the cyclopean, left and right views as CNN inputs. The two SIQA frameworks share the same implementation approach: First, to overcome the obstacle of limited SIQA datasets, we accept image patches that have been cropped from corresponding stereopairs as inputs for local quality-sensitive feature extraction. Next, a local feature selection algorithm is used to remove related features on non-salient patches, which could cause large prediction errors. Finally, the reserved local visual structures of salient regions are aggregated into a final quality score in an end-to-end manner. Experimental results on three public SIQA databases demonstrate that our method outperforms most state-of-the-art no-reference (NR) SIQA methods. The results of a cross-database experiment also show the robustness and generality of the proposed method. Guangming Sun, Bufan Shi, Andrey S. Krylov, Yong Ding 0003 |
IEEE Trans. Multim. | 5 |
| 2019 | Stereoscopic image quality assessment by analysing visual hierarchical structures and binocular effectsabstractIn the design of three‐dimensional image processing systems, stereoscopic image quality assessment (SIQA) plays an indispensable role as a performance evaluator and a supervisor, yet the study upon which remains immature due to the complexity of the human visual system (HVS). In this study, a novel SIQA method is proposed by extracting quality‐aware image features according to the properties of the hierarchical structure in the HVS. Especially, the interests of the primary and secondary visual cortex are taken into consideration, so that the image quality representation is constructed in a way both accurate and efficient. Moreover, influences caused by binocular effects including binocular rivalry and binocular visual discomfort are accounted to further improve the performance of the proposed method. The superiority of the proposed method is validated through experiments on public databases in comparison to state‐of‐the‐art works in terms of accuracy and robustness. Yong Ding 0003, Andrey S. Krylov |
IET Image Process. | 1 |
| 2017 | Efficient scheme of low-dose CT reconstruction using TV minimization with an adaptive stopping strategy and sparse dictionary learning for post-processingabstractRecently, low-dose computed tomography (CT) has become highly desirable because of the growing concern for the potential risks of excessive radiation. For low-dose CT imaging, it is a significant challenge to guarantee image quality while reducing radiation dosage. Compared with classical filtered backprojection algorithms, compressed sensing-based iterative reconstruction has achieved excellent imaging performance, but its clinical application is hindered due to its computational inefficiency. To promote low-dose CT imaging, we propose a promising reconstruction scheme which combines total-variation minimization and sparse dictionary learning to enhance the reconstruction performance, and properly schedule them with an adaptive iteration stopping strategy to boost the reconstruction speed. Experiments conducted on a digital phantom and a physical phantom demonstrate a superior performance of our method over other methods in terms of image quality and computational efficiency, which validates its potential for low-dose CT imaging. Yong Ding 0003, Tuo Hu |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2017 | Image quality assessment based on multi-feature extraction and synthesis with support vector regression
Yong Ding 0003 |
Signal Process. Image Commun. | 1 |
| 2016 | Image quality assessment method based on nonlinear feature extraction in kernel spaceabstractTo match human perception, extracting perceptual features effectively plays an important role in image quality assessment. In contrast to most existing methods that use linear transformations or models to represent images, we employ a complex mathematical expression of high dimensionality to reveal the statistical characteristics of the images. Furthermore, by introducing kernel methods to transform the linear problem into a nonlinear one, a full-reference image quality assessment method is proposed based on high-dimensional nonlinear feature extraction. Experiments on the LIVE, TID2008, and CSIQ databases demonstrate that nonlinear features offer competitive performance for image inherent quality representation and the proposed method achieves a promising performance that is consistent with human subjective evaluation. Yong Ding 0003 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2014 | Perceptual image quality assessment metric using mutual information of Gabor features
Yong Ding 0003, Xiaolang Yan, Andrey S. Krylov |
Sci. China Inf. Sci. | 1 |
| 2011 | An Efficient Compatibility-Based Test Data Compression and Its Decoder Architecture
Min-yong Wan, Yong Ding 0003, Xiaolang Yan |
J. Electron. Test. | 2 |
| 2011 | Current oscillations and low-frequency noises in GaAs MESFET channels with sidegating biasabstractLow-frequency noises (LFN) and noise-like oscillations (NLO) in GaAs metal semiconductor field effect transistor (MESFET) channel current were investigated under sidegating bias conditions. It was found that the fluctuations of the channel current were directly dependent upon the sidegating bias. As the sidegating bias decreased, the amplitudes of the oscillations would increase correspondingly. Furthermore, the LFN and NLO would attenuate sharply when the sidegating bias increased to more than a certain voltage. Two mechanisms are presented to demonstrate that the effective substrate resistivity or the channel-substrate junction modulated by sidegating bias and deep level traps would take responsibilities for the LFN and NLO. Yong Ding 0003, Xiao-hua Luo, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 1 |
| 2011 | Efficient implementation of a cubic-convolution based image scaling engineabstractIn video applications, real-time image scaling techniques are often required. In this paper, an efficient implementation of a scaling engine based on 4×4 cubic convolution is proposed. The cubic convolution has a better performance than other traditional interpolation kernels and can also be realized on hardware. The engine is designed to perform arbitrary scaling ratios with an image resolution smaller than 2560×1920 pixels and can scale up or down, in horizontal or vertical direction. It is composed of four functional units and five line buffers, which makes it more competitive than conventional architectures. A strict fixed-point strategy is applied to minimize the quantization errors of hardware realization. Experimental results show that the engine provides a better image quality and a comparatively lower hardware cost than reference implementations. Yong Ding 0003, Ming-Yu Liu 0002, Xiaolang Yan |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | A robust motion estimation with center-biased diamond search and its parallel architecture for motion-compensated de-interlace
Yong Ding 0003, Xiaolang Yan |
J. Supercomput. | 1 |