EDBT 2026 Demo / reviewers in the wild / expert
Jianguo Hu
dblp:156/2026
· DBLP profile ↗
25ranked-venue papers
0as first author
19since 2021 · last 2027
0009-0005-8872-8593ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Systems, architecture and hardware · 7 · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | A foreground-aware and consistency-guided detection transformer for industrial surface defect inspection
Jianguo Hu, Zhi-Liang Hu |
Expert Syst. Appl. | 3 |
| 2026 | GeoRoute: Signoff-Clean Analog Routing via Pin-Aligned Graph Construction for Robust Pin Access: GeoRoute: Signoff-Clean Analog Routing via Pin-Aligned Graph Construction for Robust Pin AccessabstractExisting analog layout routers that depend on a fixed global routing grid fail when pin locations do not align with the grid, a situation common in hard intellectual-property (IP) macro integration and manual designs. This brief presents GeoRoute, a geometry-driven router that constructs an instance-specific, pin-aligned graph for each connection. A multi-constraint A* search with in-search design-rule checking (DRC) and a two-stage intra-path conflict resolution mechanism efficiently produces signoff-clean paths. Across multiple process nodes and benchmarks, GeoRoute achieves 100% completion with zero DRC violations and layout-versus-schematic (LVS) clean results; reference routers average 67–92% completion. Fabrication and post-silicon validation of a 180 nm near-field communication analog front-end (NFC-AFE) further demonstrate layout quality and engineering feasibility. Jiakai Pan, Wenjia Wang 0010, Shengzhi Shen, Jianguo Hu |
ACM Great Lakes Symposium on VLSI | 6 |
| 2026 | R-SpecTTTra: Robust Song-Level AI Music Deepfake Detection under Real-World Post-Processing
Ruicheng Zou, Chengxin Chen, Nanli Zeng, Jianguo Hu |
ICIC (7) | 4 |
| 2026 | Domain-Category Fusion Guided Diffusion Model for cross-dataset facial expression recognition
Jingjie Yan, Yuebo Yue, Jinsheng Wei, Jianguo Hu |
Comput. Vis. Image Underst. | 5 |
| 2026 | Enhancing multimodal large language models with efficient feature alignment and processing using state space models
Jiakai Pan, Zhengzhuo Wang, Shengzhi Shen, Jianguo Hu |
Neurocomputing | 7 |
| 2026 | A Low-Cost 0.28 mm2 Dual-Mode ASK Demodulator NFC Forum Type 5 Compatible Fully Integrated IoT TagabstractWith the increasing prevalence of NFC-enabled smartphones, next-generation IoT devices demand contactless information interaction capabilities. This paper presents an NFC tag chip architecture for large-scale IoT applications, achieving reliable data transmission through RF coupling. Addressing the critical challenge of balancing power consumption and cost in industrial deployment, the design features a dual-mode ASK demodulator (supporting 100% and 10% modulation depth) compliant with NFC Forum Type 5 specifications, achieving ultra-low power consumption of$83.06\mu $W through dynamic clock division and gating techniques. For cost optimization, the architecture employs a byte-level anti-collision protocol for memory structure optimization and replaces conventional comparators with inverter chains. Implemented in SMIC 0.13$\mu $m EEPROM process, the compact layout of$544.24\mu \mathrm{m} \times 516.60\mu $m maintains unit cost below 1 cent. Experimental results demonstrate stable communication under maximum$9.44\mu $s modulation gap through RC-delay pulse shaping and LDO with <7% ripple. This research provides an economical and reliable hardware solution for billion-scale IoT node deployment. Yu-Xuan Huang, Jin-Biao Zhong, Hao-Cheng Hu, Ke-Xuan Chen, Yong-Jun Wen, De-Ming Wang, Jianguo Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2026 | ALAG: A Silicon-Proven, Non-Fixed-Grid Framework for Hierarchical Analog Layout Automation
Jiakai Pan, Wenjia Wang 0010, Shengzhi Shen, Zhengzhuo Wang, Jianguo Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Enhancing Visual Understanding in Multimodal Large Language Models with Efficient Feature Alignment and State Space ModelsabstractMultimodal Large Language Models (MLLMs) excel at processing complex tasks involving visual and textual data. However, existing Mamba-based MLLMs face significant challenges in visual feature extraction, resulting in poor cross-modal alignment and compromised performance. To tackle these issues, we present ML-Mamba, a novel architecture built on the Mamba framework to enhance multimodal learning. ML-Mamba features a robust visual encoder, a Mamba-Transformer Projector for optimizing feature alignment and interaction between modalities, and the advanced Mamba large language model. By utilizing techniques such as cluster-based scanning for improved visual feature extraction and a Shared-Specialized feed-forward mechanism, ML-Mamba enhances visual representation quality and overall model efficiency. Extensive benchmarking shows that ML-Mamba outperforms existing models in multimodal tasks, significantly improving inference speed and cross-modal alignment. This work underscores the potential of integrating structured state-space models with advanced transformer components to develop scalable and resource-efficient multimodal models. Jiakai Pan, Jiahao Tang, Yifei Xing 0001, Zhengzhuo Wang, Shengzhi Shen, Jianguo Hu |
ECAI | 8 |
| 2025 | VisCompConText: Scaling Multi-Modal Contexts via Visual Token Compression and Language Model GuidanceabstractTraining multimodal models with extended text contexts faces significant challenges due to prohibitive GPU memory consumption and computational costs from traditional tokenization methods, compounded by limited semantic alignment between visual and linguistic representations. Our approach introduces two key innovations: (1) Visual tagging for processing long contexts: This component adaptively renders long text segments into spatially efficient visual tokens, significantly reducing GPU memory usage and floating-point operations (FLOPs) during both training and inference. (2) LLM-Guided Visual Encoder: By leveraging large language models (LLMs), we enhance the visual encoder’s ability to comprehend long-form text, and overcome the limited long-context semantic awareness of visual encoders. Experimental results demonstrate that VisCompConText complements existing methods for extending text context length, improving document understanding, and showing strong potential in long text Q&A tasks. This work presents the first systematic solution for efficient long-context multimodal learning without sacrificing semantic granularity. Zhengzhuo Wang, Jiakai Pan, Shengzhi Shen, Jiahao Tang, Chaoxing Zhou, Jianguo Hu |
ECAI | 8 |
| 2025 | MGARoute: Efficient Analog Routing With Multi-Stage Rip-Up and Rerouting Under Geometric and Symmetry ConstraintsabstractRouting is a critical path-planning challenge in integrated circuit (IC) design, especially in analog IC design, where the complexity of the routing problem increases significantly owing to complex constraints and process parameters. Traditional automated routing techniques often rely on manual experience, limiting the effectiveness of the algorithms. To address this limitation, this study proposes a multilevel fast detour and rerouting algorithm based on reinforcement learning that integrates multiple symmetry and geometric constraints. The algorithm can effectively improve the design efficiency and reduce human error while optimizing the circuit performance and power consumption. The results show that the framework significantly improves the routing speed, reduces the number of vias, and is essentially free of DRC reporting errors on a variety of circuits in SMIC’s 180 nm technology. In addition, the algorithm roughly agrees with the manual routing results in terms of simulation performance and maintains good performance. Deming Wang, Song-Yan Jiang, Jiakai Pan, Jianguo Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | An Implementation Method for 100% ASK Modulation Applied to NFC TagsabstractPassive near field communication (NFC) tags rely on the carrier-provided clock for operation. They can receive 100% amplitude shift keying (ASK)-modulated information but are unable to respond using 100% ASK modulation. This limitation restricts the tag’s resistance to interference and its communication range. This article proposes a design approach that enables passive NFC tags to employ 100% ASK modulation (termed “strong modulation” in this article, while non-100% ASK modulation is referred to as “typical modulation”) for response. Addressing the critical issue where the tag’s clock is lost due to the antenna carrier being turned off during the low signal bits of the modulation signal, preventing the tag from continuing to function, this article introduces a high-precision recovery clock circuit as a solution. The recovery clock circuit consists of a digitally controlled oscillator (DCO) circuit composed of 12 sets of current mirrors and a logic circuit DCO calibrator. The design feasibility was validated through the postlayout parasitic extraction and the AMS mixed-signal simulation in Cadence Virtuoso, ensuring correct communication between the tag and the reader. By implementing the tag’s strong modulation response, the anti-interference capability of the tag’s returned signal can be significantly enhanced, effectively reducing the difficulty of demodulation at the receiving end and improving the tag’s poor long-distance communication capabilities. Comparatively, the minimum antenna coupling coefficientkrequired for response under strong modulation is only 40.91% of that needed for typical modulation, enabling the tag to operate in weaker electromagnetic fields and exhibit better long-distance communication capabilities. Deming Wang, Ke-Xuan Chen, Guan-Jin Xu, Jing Wu 0009, Yu-Xuan Huang, Jianguo Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | A 0.35mm2 94.25μ W Fully Integrated NFC Tag IC Using 0.13μ m CMOS ProcessabstractAn NFC tag chip employing the ISO/IEC14443-A protocol is developed in SMIC 0.13$\upmu$m EEPROM 2P6M CMOS process, boasting a chip area of 620.03$\upmu$m$\times$567.93$\upmu$m and a power consumption of under 100$\upmu$W. The chip addresses critical issues such as reducing power consumption and minimizing area costs for covering numerous IoT nodes. For a small area of the chip, a specialized ESD protection circuit is proposed, efficiently multiplexing discharge transistors, cross-gate connected rectifiers, limiters for overvoltage protection, and load switch modulators within the chip’s limited space. For power efficiency, a compact 175.44$\upmu$m$\times$32.98$\upmu$m 18.52$\upmu$A LDO based on the current mirror and current feedback is presented for the DC supply during the 100% ASK modulation. Additionally, an ASK demodulator and a 13.56MHz$\pm$kHz clock generator are designed in compact areas of 62.19$\upmu$m$\times$56.06$\upmu$m and 24.76$\upmu$m$\times$13.14$\upmu$m, respectively. These components ensure stable protocol communication with digital circuits and support a 7kb EEPROM, providing a comprehensive solution for NFC-based IoT information collection. De-Ming Wang, Jing Wu 0009, Jian-Hao Cai, Qinghua Zhong, Jianguo Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | A High Precision Analog Temperature Compensated Crystal Oscillator Using a New Temperature Compensated MultiplierabstractThe development and application of various precision electronic devices require a large number of high-precision and low-cost temperature-compensated crystal oscillators (TCXO) to generate frequency references. In order to achieve higher compensation accuracy at lower cost, a new fully integrated analog-TCXO (ATCXO) is proposed in this paper. It uses an innovative temperature-compensated multiplier to generate cubic and quintic voltages. The experimental results show that it can achieve the ultra-high fitting of the frequency-temperature characteristic curve of the crystal. This effectively improves the compensation accuracy of the TCXO. The proposed ATCXO achieves a frequency stability of ± 0.8 ppm in the temperature range of$- 40\,\,^{\circ }\text{C}$to 95 °C. At the same time, the gain-temperature error of the proposed temperature compensated analog multiplier is$\le 0.13$dB. This paper is based on SMIC 130 nm process, the core area of the chip is only 0.124 mm2. Multiple modules such as bandgap circuit, n-th power voltage generation circuit, summation circuit, gain control circuit and voltage controlled oscillator using variable capacitor array are integrated. Zi-Qi Zeng, Jianguo Hu, Jing Wu 0009, Qinghua Zhong, Deming Wang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | 3D3M: 3D Modulated Morphable Model for Monocular Face Reconstructionabstract3D face reconstruction from a single image is a vital task in various multimedia applications. A key challenge for 3D face shape reconstruction is to build the correct dense face correspondence between the monocular input face and the deformable mesh. Most existing methods rely on shape labels fitted by traditional methods or strong priors such as multi-view geometry consistency. In contrast, we propose an innovative 3D Modulated Morphable Model (3D3M) to learn the dense shape correspondence from monocular images in a self-supervised manner. Specifically, given a batch of input faces, 3D3M encodes their 3DMM attributes (shape, texture, lighting, etc.) and then randomly shuffles the 3DMM attributes to generate the attribute-changed faces. The attribute-changed faces can be encoded and rendered back in a cycle-consistent manner, which enables us to utilize the self-supervised consistencies in dense mesh vertices and reconstructed pixels. The dense shape and pixel correspondence enable us to adopt a series of self-supervised constraints to fit the 3D face model accurately and learn the per-vertex correctives end-to-end. 3D3M builds excellent high-quality 3D face reconstruction results from monocular images. Both quantitative and qualitative experimental results have verified the superiority of 3D3M over prior arts on 3D face reconstruction and face alignment. Yong Li 0032, Jianguo Hu, Xinmiao Pan, Zechao Li, Zhen Cui 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | A Fully Integrated Low-Cost HF Multistandard RFID Reader SoC and Module for IoT ApplicationsabstractThe deployment and application of large-scale Internet of Things require a large number of 13.56-MHz radio-frequency identification (RFID) reader chips with lower cost. The existing RFID reader module needs to integrate multiple chips, such as microcontroller and high-frequency (HF) RFID reader IC on a single printed circuit board (PCB), which has the bottleneck of cost, power consumption, stability, and reliability. In order to adapt to the development trend of high integration, miniaturization, and low-power consumption of RFID reader, this article proposes a fully integrated low-cost RFID reader System-on-a-Chip (SoC), which integrates RF transceiver and analog circuit, baseband protocol processing, microcontroller, memory, and interface circuit into a single chip. The proposed reader IC supports communication protocols, such as ISO/IEC 14443 Type A and Type B, ISO/IEC 15693, and ISO/IEC 7816. The reader can communicate with contact and contactless IC cards. The modulation depth of the transmitter circuit ranges from 1% to 100%, and the maximum transmitting current is 126 mA. The receiver circuit is composed of I/Q clock generation, switched capacitor sampling circuit, variable gain amplifier (VGA) with filter, comparator, and pulse shaping circuit, which has good anti-noise performance. This 13.56-MHz RFID reader SoC is fabricated in a 180-nm CMOS process and is housed in a 64-pin QFP package. The total area is 8.75 mm2, including 1.47 mm2for analog and RF circuits. Deming Wang, Jianguo Hu, Jing Wu 0009 |
IEEE Internet Things J. | 2 |
| 2021 | Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the MotionabstractOne significant factor we expect the video representation learning to capture, especially in contrast with the image representation learning, is the object motion. However, we found that in the current mainstream video datasets, some action categories are highly related with the scene where the action happens, making the model tend to degrade to a solution where only the scene information is encoded. For example, a trained model may predict a video as playing football simply because it sees the field, neglecting that the subject is dancing as a cheerleader on the field. This is against our original intention towards the video representation learning and may bring scene bias on a different dataset that can not be ignored. In order to tackle this problem, we propose to decouple the scene and the motion (DSM) with two simple operations, so that the model attention towards the motion information is better paid. Specifically, we construct a positive clip and a negative clip for each video. Compared to the original video, the positive/negative is motion-untouched/broken but scene-broken/untouched by Spatial Local Disturbance and Temporal Local Disturbance. Our objective is to pull the positive closer while pushing the negative farther to the original clip in the latent space. In this way, the impact of the scene is weakened while the temporal sensitivity of the network is further enhanced. We conduct experiments on two tasks with various backbones and different pre-training datasets, and find that our method surpass the SOTA methods with a remarkable 8.1% and 8.8% improvement towards action recognition task on the UCF101 and HMDB51 datasets respectively using the same backbone. Ke Li 0015, Jianguo Hu, Xinyang Jiang, Rongrong Ji, Xing Sun 0001 |
AAAI | 4 |
| 2021 | Self-Supervised Mutual Learning for Video Representation LearningabstractThis work tackles the problem of self-supervised learning of video representation tasks. The related works construct different surrogate supervision signals from data itself. Instead of proposing novel signal, our main insight is that the field of self-supervised learning can be benefited from mutual learning, that is, these supervision signals can learn from others and the combination between them leads to better representation. Unifying these two approaches, we present a frame-work called Self-supervised Mutual (SSM) Learning: a simple framework for mutual learning of video representation under the content of self-supervised. In order to understand what enables the task to learn useful representation, we systematically study the major components of our framework. We show that (1) surrogate supervision signal can learn from others effectively under the framework of mutual learning; (2) introducing a learnable align unit between the deep features supervised by multiple supervision signals in hidden space improves the quality of the learned representation. By combining these findings, we are able to considerably outperform previous methods for self-supervised learning on HMDB51 and UCF101 when applied to action recognition tasks. Jianguo Hu, Xuebin Yang, Yanyu Ding |
ICME | 3 |
| 2021 | Learning to locate for fine-grained image recognition
Jianguo Hu, Shiren Li |
Comput. Vis. Image Underst. | 2 |
| 2021 | Revisiting Hard Example for Action RecognitionabstractVideo-based action recognition, which needs to handle temporal motion and spatial cues simultaneously, remains a challenging task. In this paper, our motivation is to address this issue by fully utilizing temporal information. Specially, a novel light-weight Voting-based Temporal Correlation (VTC) module is proposed to enhance temporal information. Multiple branches with different temporal sampling intervals are included in this module and they are regarded as voters. The final classification result is “voted” by these branches together. VTC module integrates sparse temporal sampling strategy into feature sequences, so it mitigates the effect of redundant information and focuses more on temporal modeling. Additionally, we propose a simple and intuitive Similarity Loss (SL) to guide the training procedure of the VTC module and the backbone network. When we introduce confusion in the predicted vector intentionally, SL eases intra-class variation by discovering class-specific common motion patterns rather than sample-specific discriminative information. SL neither needs excessive parameter tuning during training nor adds significant computation overhead during test time. By combining VTC module and SL with complementary advances in the field, we clearly outperform state-of-the-art results and achieve 83.0, 98.4, 49.6 and 77.8 accuracy on HMDB51, UCF101, something-something-v1, and Kinetics respectively. Jianguo Hu, Shiren Li, Zhihao Yuan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Combination of temporal-channels correlation information and bilinear feature for action recognitionabstractIn this study, the authors focus on improving the spatio–temporal representation ability of three‐dimensional (3D) convolutional neural networks (CNNs) in the video domain. They observe two unfavourable issues: (i) the convolutional filters only dedicate to learning local representation along input channels. Also they treat channel‐wise features equally, without emphasising the important features; (ii) traditional global average pooling layer only captures first‐order statistics, ignoring finer detail features useful for classification. To mitigate these problems, they proposed two modules to boost 3D CNNs’ performance, which are temporal‐channel correlation (TCC) and bilinear pooling module. The TCC module can capture the information of inter‐channel correlations over the temporal domain. Moreover, the TCC module generates channel‐wise dependencies, which can adaptively re‐weight the channel‐wise features. Therefore, the network can focus on learning important features. With regards to the bilinear pooling module, it can capture more complex second‐order statistics in deep features and generate a second‐order classification vector. We can get more accurate classification results by combining the first‐order and second‐order classification vector. Extensive experiments show that adding our proposed modules to I3D network could consistently improve the performance and outperform the state‐of‐the‐art methods. The code and models are available at https://github.com/caijh33/I3D_TCC_Bilinear . Jiahui Cai, Jianguo Hu, Shiren Li, Jialing Lin |
IET Comput. Vis. | 2 |
| 2020 | 3D RANs: 3D Residual Attention Networks for action recognition
Jiahui Cai, Jianguo Hu |
Vis. Comput. | 2 |
| 2018 | Pose Transferrable Person Re-IdentificationabstractPerson re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To address this issue, we propose a pose-transferrable person ReID framework which utilizes pose-transferred sample augmentations (i.e., with ID supervision) to enhance ReID model training. On one hand, novel training samples with rich pose variations are generated via transferring pose instances from MARS dataset, and they are added into the target dataset to facilitate robust training. On the other hand, in addition to the conventional discriminator of GAN (i.e., to distinguish between REAL/FAKE samples), we propose a novel guider sub-network which encourages the generated sample (i.e., with novel pose) towards better satisfying the ReID loss (i.e., cross-entropy ReID loss, triplet ReID loss). In the meantime, an alternative optimization procedure is proposed to train the proposed Generator-Guider-Discriminator network. Experimental results on Market-1501, DukeMTMC-reID and CUHK03 show that our method achieves great performance improvement, and outperforms most state-of-the-art methods without elaborate designing the ReID model. Jinxian Liu, Bingbing Ni, Yichao Yan, Peng Zhou 0010, Jianguo Hu |
CVPR | 6 |
| 2018 | Crowd Counting via Adversarial Cross-Scale Consistency PursuitabstractCrowd counting or density estimation is a challenging task in computer vision due to large scale variations, perspective distortions and serious occlusions, etc. Existing methods generally suffer from two issues: 1) the model averaging effects in multi-scale CNNs induced by the widely adopted ℓ2regression loss; and 2) inconsistent estimation across different scaled inputs. To explicitly address these issues, we propose a novel crowd counting (density estimation) framework called Adversarial Cross-Scale Consistency Pursuit (ACSCP). On one hand, a U-net structured generation network is designed to generate density map from input patch, and an adversarial loss is directly employed to shrink the solution onto a realistic subspace, thus attenuating the blurry effects of density map estimation. On the other hand, we design a novel scale-consistency regularizer which enforces that the sum up of the crowd counts from local patches (i.e., small scale) is coherent with the overall count of their region union (i.e., large scale). The above losses are integrated via a joint training scheme, so as to help boost density estimation performance by further exploring the collaboration between both objectives. Extensive experiments on four benchmarks have well demonstrated the effectiveness of the proposed innovations as well as the superior performance over prior art. Zan Shen, Yi Xu 0001, Bingbing Ni, Minsi Wang, Jianguo Hu, Xiaokang Yang 0001 |
CVPR | 5 |
| 2018 | Scale-Transferrable Object DetectionabstractScale problem lies in the heart of object detection. In this work, we develop a novel Scale-Transferrable Detection Network (STDN) for detecting multi-scale objects in images. In contrast to previous methods that simply combine object predictions from multiple feature maps from different network depths, the proposed network is equipped with embedded super-resolution layers (named as scale-transfer layer/module in this work) to explicitly explore the interscale consistency nature across multiple detection scales. Scale-transfer module naturally fits the base network with little computational cost. This module is further integrated with a dense convolutional network (DenseNet) to yield a one-stage object detector. We evaluate our proposed architecture on PASCAL VOC 2007 and MS COCO benchmark tasks and STDN obtains significant improvements over the comparable state-of-the-art detection models. Peng Zhou 0010, Bingbing Ni, Cong Geng, Jianguo Hu, Yi Xu 0001 |
CVPR | 4 |
| 2018 | Live Face Verification with Multiple Instantialized Local Homographic ParameterizationabstractState-of-the-art live face verification methods would easily be attacked by recorded facial expression sequence. This work directly addresses this issue via proposing a patch-wise motion parameterization based verification network infrastructure. This method directly explores the underlying subtle motion difference between the facial movements re-captured from a planer screen (e.g., a pad) and those from a real face; therefore interactive facial expression is no longer required. Furthermore, inspired by the fact that ?a fake facial movement sequence MUST contains many patch-wise fake sequences?, we embed our network into a multiple instance learning framework, which further enhance the recall rate of the proposed technique. Extensive experimental results on several face benchmarks well demonstrate the superior performance of our method. Chen Lin 0001, Zhouyingcheng Liao, Peng Zhou 0010, Jianguo Hu, Bingbing Ni |
IJCAI | 4 |