EDBT 2026 Demo / reviewers in the wild / expert
Hong Zhang 0009
dblp:24/6914-9
· DBLP profile ↗
15ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 since 2021Systems, architecture and hardware · 7 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D Lane Detection Based on Projection-Consistent Reference Points and Intra- & Inter-lane Contextabstract3D lane detection aims to identify lane categories and trends in 3D space, which is a vital and challenging task in autonomous driving. Existing methods introduce various priors to guide 3D lane prediction, which generally consist of a series of reference points for context aggregation. However, due to the misalignment between these reference points and the lanes, it is difficult to obtain complete and discriminative context for complex instances. In this paper, we are devoted to introducing 3D priors adaptive to lane appearances, which serve as references to aggregate the lane context. Specifically, we propose a projection-consistent reference generation strategy to keep the projected 3D reference points geometrically consistent with the corresponding lanes in images. In addition, a segmentation-lifting denoising strategy is designed to improve the ability of the model to map the lane segmentation into 3D space. To leverage more lane-related information, we propose a decoupled lane-context aggregation module by considering the perspectives of individual geometries and integrated layout, namely intra-lane and inter-lane context. Extensive experiments on the OpenLane dataset show that our approach outperforms previous methods and achieves the state-of-the-art performance. The code will be made publicly available. Yiqiu Bing, Huilin Niu, Hong Zhang 0009, Zhong Zhou, Qichuan Geng |
ICRA | 3 |
| 2024 | Artificial Neural Network Based Calibration for a 12 b 250 MS/s Pipelined-SAR ADC With Ring Amplifier in 40-nm CMOSabstractThis paper presents a 2-stage pipelined-SAR ADC with artificial-neural-network (ANN) based digital calibration algorithm to calibrate the mismatch error in the$1^{\mathrm {st}}$-stage capacitive DAC (CDAC) and the inter-stage gain error (IGE) together. Previous ANN-based calibration schemes suffer from excessive power and hardware overhead due to the large number of network parameters. To facilitate hardware implementation, the proposed algorithm only requires$N_{1}+1$input parameters ($N_{1}$is the resolution of the$1^{\mathrm {st}}$-stage SAR ADC), in which the overall output of the$2^{\mathrm {nd}}$-stage SAR ADC is combined into a single parameter. In addition, the ANN utilizes a single-neuron hidden layer with linear activation function to calculate the actual bit weight of the ADC, remarkably reducing the hardware overhead and power consumption of the calibration circuit. The prototype 12-bit, 250 MS/s pipelined-SAR ADC with “loop-unrolled” architecture is implemented in 40-nm CMOS, in which a ring amplifier with improved bias circuit is used to realize a robust closed-loop gain for residue amplification. With the ANN-based calibration circuit implemented in an FPGA, the calibrated ADC achieves the SNDR of 65.0 dB and the SFDR of 84.0 dB at Nyquist input (124 MHz), with a Schreier figure of merit of 169.0 dB and a Walden figure of merit of 14.0 fJ/conv-step. The ADC core consumes 4.95 mW, with an active area of only 0.013 mm2. Bin Liu 0068, Zhichao Dai, Yufeng Ge, Huanhuan Qi, Jie Zhang 0039, Zhenhai Chen, Yan Xue, Hong Zhang 0009 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 13 |
| 2022 | Part-Level Car Parsing and Reconstruction in Single Street View ImagesabstractPart information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel. Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | A 2.5-MHz BW, 75-dB SNDR Noise-Shaping SAR ADC With a 1st-Order Hybrid EF-CIFF Structure Assisted by Unity-Gain BufferabstractThis article presents a 1st-order noise-shaping (NS) successive approximation register (SAR) analog-to-digital converter (ADC) with a hybrid error-feedback (EF) and cascaded-integrator-feed-forward (CIFF) structure assisted by a unity-gain buffer (UGB). Without using a multi-input comparator which is widely adopted in conventional 1st-order passive NS structures, the proposed hybrid EF-CIFF structure realizes a more ideal 1st-order noise transfer function (NTF) with a reasonable capacitance ratio, so as to obtain better NS effect. Fabricated in a 28-nm CMOS technology, the prototype NS-SAR ADC consumes 150$\mu \text{W}$under a 0.9-V supply voltage when operating at 40-MS/s sampling rate. A 75-dB signal-to-noise-and-distortion ratio (SNDR) is measured for a 2.47-MHz sinusoid input under an oversampling ratio (OSR) of 8. It achieves a peak Schreier figure-of-merit (FoM) of 177.2 dB and the core circuit occupies 0.012-mm2 area. Hanrui Zhang 0007, Zihao Jiao, Di Mu, Jie Zhang 0039, Hong Zhang 0009 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2021 | Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingabstractImage-to-image translation has been revolutionized with GAN-based methods. However, existing methods lack the ability to preserve the identity of the source domain. As a result, synthesized images can often over-adapt to the reference domain, losing important structural characteristics and suffering from suboptimal visual quality. To solve these challenges, we propose a novel frequency domain image translation (FDIT) framework, exploiting frequency information for enhancing the image generation process. Our key idea is to decompose the image into low-frequency and high-frequency components, where the high-frequency feature captures object structure akin to the identity. Our training objective facilitates the preservation of frequency information in both pixel space and Fourier spectral space. We broadly evaluate FDIT across five large-scale datasets and multiple tasks including image translation and GAN inversion. Extensive experiments and ablations show that FDIT effectively preserves the identity of the source image, and produces photo-realistic images. FDIT establishes state-of-the-art performance, reducing the average FID score by 5.6% compared to the previous best method. Mu Cai, Hong Zhang 0009, Huijuan Huang 0001, Qichuan Geng, Yixuan Li 0001, Gao Huang 0001 |
ICCV | 2 |
| 2021 | A 1st-Order Passive Noise-Shaping SAR ADC with Improved NTF Assisted by Comparator Gain CalibrationabstractThis paper presents an improved 1st-order fully passive noise-shaping (NS) scheme for NS SAR ADCs. In order to move the zero and pole of the noise transfer function (NTF) closer to the unit circle, a structure combining charge pump and two-input comparator is used to provide two gain factors of 2 and 3, respectively, forming an NTF with zero and pole located at (0.875, 0) and (-0.75, 0), respectively. The improved NTF can provide better noise suppression effect than conventional NTFs. To ensure the stability of the ADC, a gain calibration technique is proposed for the two-input comparator to ensure the gain ratio between the 2 input branches, which can prevent the pole from moving outside the unit circle because of the gain mismatch between the comparator's 2 input branches. In addition, the proposed structure can merge the integration cycle into the sampling phase to enhance the conversion speed. Based on the proposed idea, an NS SAR ADC with a 10-bit DAC is designed in 65-nm CMOS, with post-layout simulation results showing that a 76-dB SNDR is achieved with 2.5-MHz signal bandwidth and 40-MS/s sampling rate. Hanrui Zhang 0007, Zihao Jiao, Jie Zhang 0039, Hong Zhang 0009 |
ISCAS | 6 |
| 2021 | Image Re-composition via Regional Content-Style DecouplingabstractTypical image composition harmonizes regions from different images to a single plausible image. We extend the idea of image composition by introducing the content-style decomposition and combination to form the concept of image re-composition. In other words, our image re-composition could arbitrarily combine those contents and styles decomposed from different images to generate more diverse images in a unified framework. In the decomposition stage, we incorporate the whitening normalization to obtain a more thorough content-style decoupling, which substantially improves the re-composition results. Moreover, to handle the variation of structure and texture of different objects in an image, we design the network to support regional feature representation and achieve region-aware content-style decomposition. Regarding the composition stage, we propose a cycle consistency loss to constrain the network preserving the content and style information during the composition. Our method can produce diverse re-composition results, including content-content, content-style and style-style. Our experimental results demonstrate a large improvement over the current state-of-the-art methods. Wei Li 0111, Hong Zhang 0009, Ruigang Yang, Weiwei Xu 0003 |
ACM Multimedia | 4 |
| 2021 | Gated Path Selection Network for Semantic SegmentationabstractSemantic segmentation is a challenging task that needs to handle large scale variations, deformations, and different viewpoints. In this paper, we develop a novel network named Gated Path Selection Network (GPSNet), which aims to adaptively select receptive fields while maintaining the dense sampling capability. In GPSNet, we first design a two-dimensional SuperNet, which densely incorporates features from growing receptive fields. And then, a Comparative Feature Aggregation (CFA) module is introduced to dynamically aggregate discriminative semantic context. In contrast to previous works that focus on optimizing sparse sampling locations on regular grids, GPSNet can adaptively harvest free form dense semantic context information. The derived adaptive receptive fields and dense sampling locations are data-dependent and flexible which can model various contexts of objects. On two representative semantic segmentation datasets, i.e., Cityscapes and ADE20K, we show that the proposed approach consistently outperforms previous methods without bells and whistles. Qichuan Geng, Hong Zhang 0009, Xiaojuan Qi 0001, Gao Huang 0001, Ruigang Yang, Zhong Zhou |
IEEE Trans. Image Process. | 2 |
| 2019 | Improved Techniques for Training Adaptive Deep NetworksabstractAdaptive inference is a promising technique to improve the computational efficiency of deep models at test time. In contrast to static models which use the same computation graph for all instances, adaptive networks can dynamically adjust their structure conditioned on each input. While existing research on adaptive inference mainly focuses on designing more advanced architectures, this paper investigates how to train such networks more effectively. Specifically, we consider a typical adaptive deep network with multiple intermediate classifiers. We present three techniques to improve its training efficacy from two aspects: 1) a Gradient Equilibrium algorithm to resolve the conflict of learning of different classifiers; 2) an Inline Subnetwork Collaboration approach and a One-for-all Knowledge Distillation algorithm to enhance the collaboration among classifiers. On multiple datasets (CIFAR-10, CIFAR-100 and ImageNet), we show that the proposed approach consistently leads to further improved efficiency on top of state-of-the-art adaptive deep networks. Hao Li 0069, Hong Zhang 0009, Xiaojuan Qi 0001, Ruigang Yang, Gao Huang 0001 |
ICCV | 2 |
| 2019 | Implicit Semantic Data Augmentation for Deep NetworksabstractIn this paper, we propose a novel implicit semantic data augmentation (ISDA) approach to complement traditional augmentation techniques like flipping, translation or rotation. Our work is motivated by the intriguing property that deep networks are surprisingly good at linearizing features, such that certain directions in the deep feature space correspond to meaningful semantic transformations, e.g., adding sunglasses or changing backgrounds. As a consequence, translating training samples along many semantic directions in the feature space can effectively augment the dataset to improve generalization. To implement this idea effectively and efficiently, we first perform an online estimate of the covariance matrix of deep features for each class, which captures the intra-class semantic variations. Then random vectors are drawn from a zero-mean normal distribution with the estimated covariance to augment the training data in that class. Importantly, instead of augmenting the samples explicitly, we can directly minimize an upper bound of the expected cross-entropy (CE) loss on the augmented training set, leading to a highly efficient algorithm. In fact, we show that the proposed ISDA amounts to minimizing a novel robust CE loss, which adds negligible extra computational cost to a normal training procedure. Although being simple, ISDA consistently improves the generalization performance of popular deep models (ResNets and DenseNets) on a variety of datasets, e.g., CIFAR-10, CIFAR-100 and ImageNet. Code for reproducing our results are available at https://github.com/blackfeather-wang/ISDA-for-Deep-Networks. Yulin Wang 0002, Xuran Pan, Shiji Song, Hong Zhang 0009, Gao Huang 0001, Cheng Wu 0002 |
NeurIPS | 4 |
| 2018 | A 10-bit 200-kS/s 1.76-μW SAR ADC with Hybrid CAP-MOS DAC for Energy-Limited ApplicationsabstractThis paper presents a low-power and area efficient 10-bit SAR ADC with hybrid capacitive-MOS consisting of a 7-bit MSB capacitive DAC (CDAC) and a 3-bit LSB MOS DAC, which consumes less power and much smaller chip area than a pure CDAC. Instead of using a string of 8 MOS transistors to control one unit capacitor, the 3-bit LSB MOS DAC is realized by a MOS string with 4 native MOS transistors to control 2 unit capacitors, which allows higher voltage drop and more reliable operation for each unit MOS. The overall energy consumption of the proposed CAP-MOS DAC is reduced by 56.2% compared with a Vcm-based 10-bit pure CDAC. Under a 200-kS/s conversion rate, the prototype 10-bit SAR ADC is implemented in a 0.18-μm CMOS technology, showing an SNDR/ SFDR of 56.91 dB/68.56 dB at 99-kHz input under a 0.6-V power-supply, while consuming 1.76 μW at 200 kS/s for a FoM of 15.38 fJ/step. The peak DNL and INL are +0.27/-0.21 LSB and +0.43/-0.45 LSB, respectively. The ADC occupies a small active area of 0.097 mm2. Hongshuai Zhang, Hong Zhang 0009, Ruizhi Zhang 0002 |
ISCAS | 2 |
| 2018 | A Low-Power Pipelined-SAR ADC Using Boosted Bucket-Brigade Device for Residue Charge Processing
Hong Zhang 0009, Junqiang Sun, Jie Zhang 0039, Ruizhi Zhang 0002, Anthony Chan Carusone |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Unsupervised Learning of Stereo MatchingabstractConvolutional neural networks showed the ability in stereo matching cost learning. Recent approaches learned parameters from public datasets that have ground truth disparity maps. Due to the difficulty of labeling ground truth depth, usable data for system training is rather limited, making it difficult to apply the system to real applications. In this paper, we present a framework for learning stereo matching costs without human supervision. Our method updates network parameters in an iterative manner. It starts with a randomly initialized network. Left-right check is adopted to guide the training. Suitable matching is then picked and used as training data in following iterations. Our system finally converges to a stable state and performs even comparably with other supervised methods. Chao Zhou 0001, Hong Zhang 0009, Xiaoyong Shen, Jiaya Jia |
ICCV | 2 |
| 2017 | All-Digital Calibration of Timing Mismatch Error in Time-Interleaved Analog-to-Digital ConvertersabstractThis paper presents an all-digital background calibration for timing mismatch in time-interleaved analog-to-digital converters (TI-ADCs). It combines digital adaptive timing mismatch estimation and digital derivative-based correction, achieving lower hardware cost and better suppression of timing mismatch tones than previous work. In addition, for the first time closed-form exact expressions for the signal-to-noise and distortion ratio (SNDR) of a four-channel TI-ADC with timing mismatch after derivative-based digital correction are obtained, which can be used to guide the design. Simulation results of a four-channel TI-ADC behavioral model and measurement results from a commercial 12-bit 3.6-GS/s two-channel TI-ADC show that the proposed all-digital calibration can accurately estimate the timing skew and effectively correct the timing mismatch errors, while also confirming the analytic SNDR expressions. Luke Wang, Hong Zhang 0009, Rosanah Murugesu, Dustin Dunwell, Anthony Chan Carusone |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Multi-scale Patch Aggregation (MPA) for Simultaneous Detection and SegmentationabstractAiming at simultaneous detection and segmentation (SD-S), we propose a proposal-free framework, which detect and segment object instances via mid-level patches. We design a unified trainable network on patches, which is followed by a fast and effective patch aggregation algorithm to infer object instances. Our method benefits from end-to-end training. Without object proposal generation, computation time can also be reduced. In experiments, our method yields results 62.1% and 61.8% in terms of mAPr on VOC2012 segmentation val and VOC2012 SDS val, which are state-of-the-art at the time of submission. We also report results on Microsoft COCO test-std/test-dev dataset in this paper. Shu Liu 0005, Xiaojuan Qi 0001, Jianping Shi, Hong Zhang 0009, Jiaya Jia |
CVPR | 4 |