VLDB 2026 Research / reviewers in the wild / expert
Yiqing Hu
dblp:139/3141
· DBLP profile ↗
15ranked-venue papers
3as first author
8since 2021 · last 2024
0009-0000-0342-0834ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Computer networks · 3Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enhancing Graph Random Walk Acceleration via Efficient Dataflow and Hybrid Memory ArchitectureabstractGraph random walk sampling is becoming increasingly important with the widespread popularity of graph applications. It aims to capture the desirable graph properties by launching multiple walkers to collect feature paths. However, previous research suffers long sampling latency and severe memory access bottlenecks due to intrinsic data dependency and skewed vertex distribution. Thus, in this paper, we propose FastRW, a dedicated accelerator to boost graph random walk operation on FPGAs. Specifically, FastRW first integrates multiple parallel processing engines to achieve data-level parallelism, where each processing engine also leverages dataflow scheduling to resolve data dependency and hide long sampling latency. Secondly, FastRW leverages a combination of multiple storage resources to implement a hybrid memory architecture adapted to skewed vertex distribution. By integrating the above optimizations, FastRW develops a performance model to take advantage of the balance between computation parallelism and bandwidth demand. We evaluate FastRW with two classic sampling algorithms on a wide range of real-world graph datasets. The experimental results show that FastRW achieves a speedup of 37.52$\boldsymbol{\times}$on average over the system running on two 8-core Intel CPUs. FastRW also achieves an average of 28.04$\boldsymbol{\times}$speedup over the architecture implemented on V100 GPU. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Yiqing Hu, Zhongming Liu, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 5 |
| 2023 | The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-trainingabstractThe self-supervised Masked Image Modeling (MIM) schema, following "mask-and-reconstruct" pipeline of recovering contents from masked image, has recently captured the increasing interest in the community, owing to the excellent ability of learning visual representation from unlabeled data. Aiming at learning representations with high semantics abstracted, a group of works attempts to reconstruct non-semantic pixels with large-ratio masking strategy, which may suffer from "over-smoothing" problem, while others directly infuse semantics into targets in off-line way requiring extra data. Different from them, we shift the perspective to the Fourier domain which naturally has global perspective and present a new Masked Image Modeling (MIM), termed Geminated Gestalt Autoencoder (Ge^2-AE) for visual pre-training. Specifically, we equip our model with geminated decoders in charge of reconstructing image contents from both pixel and frequency space, where each other serves as not only the complementation but also the reciprocal constraints. Through this way, more robust representations can be learned in the pre-trained encoders, of which the effectiveness is confirmed by the juxtaposing experimental results on downstream recognition tasks. We also conduct several quantitative and qualitative experiments to investigate the learning behavior of our method. To our best knowledge, this is the first MIM work to solve the visual pre-training through the lens of frequency domain. Hao Liu 0003, Xinghua Jiang, Xin Li 0118, Antai Guo, Yiqing Hu, Deqiang Jiang, Bo Ren 0002 |
AAAI | 5 |
| 2023 | TaCo: Textual Attribute Recognition via Contrastive LearningabstractAs textual attributes like font are core design elements of document format and page style, automatic attributes recognition favor comprehensive practical applications. Existing approaches already yield satisfactory performance in differentiating disparate attributes, but they still suffer in distinguishing similar attributes with only subtle difference. Moreover, their performance drop severely in real-world scenarios where unexpected and obvious imaging distortions appear. In this paper, we aim to tackle these problems by proposing TaCo, a contrastive framework for textual attribute recognition tailored toward the most common document scenes. Specifically, TaCo leverages contrastive learning to dispel the ambiguity trap arising from vague and open-ended attributes. To realize this goal, we design the learning paradigm from three perspectives: 1) generating attribute views, 2) extracting subtle but crucial details, and 3) exploiting valued view pairs for learning, to fully unlock the pre-training potential. Extensive experiments show that TaCo surpasses the supervised counterparts and advances the state-of-the-art remarkably on multiple attribute recognition tasks. Online services of TaCo will be made available. Chang Nie, Yiqing Hu, Yanqiu Qu, Hao Liu 0003, Deqiang Jiang, Bo Ren 0002 |
AAAI | 2 |
| 2022 | TDv2: A Novel Tree-Structured Decoder for Offline Mathematical Expression RecognitionabstractIn recent years, tree decoders become more popular than LaTeX string decoders in the field of handwritten mathematical expression recognition (HMER) as they can capture the hierarchical tree structure of mathematical expressions. However previous tree decoders converted the tree structure labels into a fixed and ordered sequence, which could not make full use of the diversified expression of tree labels. In this study, we propose a novel tree decoder (TDv2) to fully utilize the tree structure labels. Compared with previous tree decoders, this new model does not require a fixed priority for different branches of a node during training and inference, which can effectively improve the model generalization capability. The input and output of the model make full use of the tree structure label, so that there is no need to find the parent node in the decoding process, which simplifies the decoding process and adds a prior information to help predict the node. We verified the effectiveness of each part of the model through comprehensive ablation experiments and attention visualization analysis. On the authoritative CROHME 14/16/19 datasets, our method achieves the state-of-the-art results. Changjie Wu, Jun Du 0002, Jianshu Zhang 0001, Bo Ren 0002, Yiqing Hu |
AAAI | 7 |
| 2022 | Query-driven Generative Network for Document Information Extraction in the WildabstractThis paper focuses on solving Document Information Extraction (DIE) in the wild problem, which is rarely explored before. In contrast to existing studies mainly tailored for document cases in known templates with predefined layouts and keys under the ideal input without OCR errors involved, we aim to build up a more practical DIE paradigm for real-world scenarios where input document images may contain unknown layouts and keys in the scenes of the problematic OCR results. To achieve this goal, we propose a novel architecture, termed Query-driven Generative Network (QGN), which is equipped with two consecutive modules, i.e., Layout Context-aware Module (LCM) and Structured Generation Module (SGM). Given a document image with unseen layouts and fields, the former LCM yields the value prefix candidates serving as the query prompts for the SGM to generate the final key-value pairs even with OCR noise. To further investigate the potential of our method, we create a new large-scale dataset, named LArge-scale STructured Documents (LastDoc4000), containing 4,000 documents with 1,511 layouts and 3,500 different keys. In experiments, we demonstrate that our QGN consistently achieves the best F1-score on the new LastDoc4000 dataset by at most 30.32% absolute improvement. A more comprehensive experimental analysis and experiments on other public benchmarks also verify the effectiveness and robustness of our proposed method for the wild DIE task. Haoyu Cao 0001, Xin Li 0118, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu 0003, Yinsong Liu, Bo Ren 0002 |
ACM Multimedia | 6 |
| 2022 | Relational Representation Learning in Visually-Rich DocumentsabstractRelational understanding is critical for a number of visually-rich documents (VRDs) understanding tasks. Through multi-modal pre-training, recent studies provide comprehensive contextual representations and exploit them as prior knowledge for downstream tasks. In spite of their impressive results, we observe that the widespread relational hints (e.g., relation of key/value fields on receipts) built upon contextual knowledge are not excavated yet. To mitigate this gap, we propose DocReL, a Document Relational Representation Learning framework. The major challenge of DocReL roots in the variety of relations. From the simplest pairwise relation to the complex global structure, it is infeasible to conduct supervised training due to the definition of relation varies and even conflicts in different tasks. To deal with the unpredictable definition of relations, we propose a novel contrastive learning task named Relational Consistency Modeling (RCM), which harnesses the fact that existing relations should be consistent in differently augmented positive views. RCM provides relational representations which are more compatible to the urgent need of downstream tasks, even without any knowledge about the exact definition of relation. DocReL achieves better performance on a wide variety of VRD relational understanding tasks, including table structure recognition, key information extraction and reading order detection. Xin Li 0118, Yiqing Hu, Haoyu Cao 0001, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
ACM Multimedia | 3 |
| 2022 | GMN: Generative Multi-modal Network for Practical Document Information ExtractionabstractHaoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, Bo Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Haoyu Cao 0001, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
NAACL-HLT | 4 |
| 2021 | RecycleNet: An Overlapped Text Instance Recovery ApproachabstractText recognition is the key pillar for many real-world multimedia applications. Existing text recognition approaches focus on recognizing isolated instances, whose text fields are visually separated and have no interference with each other. Moreover, these approaches cannot handle overlapped instances that often appear in sheets like invoices, receipts and math exercises, where printed templates are generated beforehand and extra contents are added afterward on existing texts. In this paper, we aim to tackle this problem by proposing RecycleNet, which automatically extracts and reconstructs overlapped instances by fully recycling the intersecting pixels that used to be obstacles for recognition. RecycleNet parallels to existing recognition systems, and serves as a plug-and-play module to boost recognition performance with zero-effort. We also released an OverlapText-500 dataset, which helps to boost the design of better overlapped text recovery and recognition solutions. Yiqing Hu, Xinghua Jiang, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002, Rongrong Ji |
ACM Multimedia | 1 |
| 2020 | Accurate Structured-Text Spotting for Arithmetical Exercise CorrectionabstractCorrecting arithmetical exercise is a labor intensive and time consuming task for primary school teachers all the time. To reduce their burdens, we propose Arithmetical Exercise Checker (AEC), which is the first system that automatically evaluates all arithmetical expressions (AEs) on exercise images. The major challenge is that AE is formed by printed and handwritten texts with particular arithmetical patterns (e.g., multi-line, fraction). Despite being part of AE, handwritten texts usually lead to zigzag boundaries and tangled rows. What's worse, AE may be arithmetical incorrect, which makes the contextual information less valuable for recognition. To tackle these problems, we introduce integrated detection, recognition and evaluation branches by leveraging AE's intrinsic features, namely 1) boundary indistinctive, 2) locally relevant patterns and 3) globally irrelevant symbols. Experimental results demonstrate that AEC yields a 93.72% correction accuracy on 40 kinds of mainstream primary arithmetical exercises. So far, the online service of AEC processes 75, 000 arbitrary exercises on average per day, and already reduced the burden of over 1, 000, 000 users. AEC shows the benefits for implementing an vision-based system as a way to aid teachers in reducing reduplicative tasks. Yiqing Hu, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002 |
AAAI | 1 |
| 2017 | Stride-in-the-Loop Relative Positioning Between Users and Dummy Acoustic SpeakersabstractWe propose and implement a novel positioning system, WalkieLokie, which directly calculates the relative position from a smart device to a target. The requirement of the target is simple: it is attached with a “dummy” acoustic speaker, which does not have any other rich capabilities, such as audio recording, communication, or computation. Hence, the proliferation of smart devices, together with the cheap accessory (e.g., dummy speaker) embedded in daily used items (e.g., smart clothes), paves the way for WalkieLokie applications. WalkieLokie leverages the walking motion for locating an acoustic speaker. The key insight is that the distance between the user and the speaker varies in real time when the user walks, and the pattern of the variance implies the relative position. We design a novel algorithm to estimate the position and signal processing methods to support accurate positioning. The experiment results show that the mean errors of ranging and direction estimation are 0.63 m and 2.46°, respectively. Extensive experiments conducted in noisy environments validate the robustness of WalkieLokie. Wenchao Huang 0001, Xiang-Yang Li 0001, Yan Xiong 0001, Panlong Yang, Yiqing Hu, Xufei Mao, Fuyou Miao 0001, Baohua Zhao, Ju-Min Zhao |
IEEE J. Sel. Areas Commun. | 5 |
| 2016 | WalkieLokie: sensing relative positions of surrounding presenters by acoustic signalsabstractIn this paper, we propose and implement WalkieLokie, a novel acoustic-based relative positioning system. WalkieLokie facilitates a multitude of Augmented Reality (AR) applications: users with smart devices can passively acquire surrounding information in real time, similar to the commercial AR system Wikitude; the surrounding presenters, who want to share information or introduce themselves, can actively launch the function on demand. The key rational of WalkieLokie is that a user can perceive a series of spatial-related acoustic signals emitted from a presenter, which depicts the relation position between the user and the presenter. The proliferation of smart devices, together with the cheap accessory (e.g., dummy speaker) embedded in daily used items (e.g., smart clothes), paves the way for WalkieLokie applications. We design a novel algorithm to estimate the position and signal processing methods to support accurate positioning. The experiment results show that the mean error of ranging and direction estimation is 0.63m and 2.46 degrees respectively. Extensive experiments conducted in noisy environments validate the robustness of WalkieLokie. Wenchao Huang 0001, Xiang-Yang Li 0001, Yan Xiong 0001, Panlong Yang, Yiqing Hu, Xufei Mao, Fuyou Miao 0001, Baohua Zhao, Ju-Min Zhao |
UbiComp | 5 |
| 2016 | User-Demand-Oriented Privacy-Preservation in Video DeliveringabstractThis paper presents a framework for privacy-preserving video delivery system to fulfill users' privacy demands. The proposed framework leverages the inference channels in sensitive behavior prediction and object tracking in a video surveillance system for the sequence privacy protection. For such a goal, we need to capture different pieces of evidence which are used to infer the identity. The temporal, spatial and context features are extracted from the surveillance video as the observations to perceive the privacy demands and their correlations. Taking advantage of quantifying various evidence and utility, we let users subscribe videos with a viewer-dependent pattern. We implement a prototype system for off-line and on-line requirements in two typical monitoring scenarios to construct extensive experiments. The evaluation results show that our system can efficiently satisfy users' privacy demands while saving over 25% more video information compared to traditional video privacy protection schemes. Haohua Du, Taeho Jung, Xuesi Jian, Yiqing Hu, Jiahui Hou, Xiang-Yang Li 0001 |
MSN | 4 |
| 2015 | Fast Similarity Search of Multi-Dimensional Time Series via Segment Rotation
Xudong Gong, Yan Xiong 0001, Wenchao Huang 0001, Lei Chen 0002, Qiwei Lu, Yiqing Hu |
DASFAA (1) | 6 |
| 2015 | Lightitude: Indoor Positioning Using Ubiquitous Visible Lights and COTS DevicesabstractIn this paper, we propose a novel indoor localization scheme, Lightitude, by exploiting ubiquitous visible lights, which are necessarily and densely deployed in almost all indoor environments. Different from existing positioning systems that exploit special LEDs, ubiquitous visible lights lack fingerprints that can uniquely identify the light source, which results in an ambiguity problem that an RLS may correspond to multiple candidate positions. Moreover, received light strength (RLS) is not only determined by device's position, but also seriously affected by its orientation, which causes great complexity in site-survey. To address these challenges, we first propose and validate a realistic light strength model to avoid the expensive site-survey, then harness user's mobility to generate spatial-related RLS to tackle single RLS's position-ambiguity problem. Experiment results show that Lightitude achieves mean accuracy 1.93m and 2.24m in office (720m2) and library scenario (960m2) respectively. Yiqing Hu, Yan Xiong 0001, Wenchao Huang 0001, Xiang-Yang Li 0001, Xufei Mao, Panlong Yang, Caimei Wang |
ICDCS | 1 |
| 2015 | Magemite: Character inputting system based on magnetic sensorabstractWe propose Magemite, a fine-grained input system that exploits the around device space (ADS) as an expansion of the limited input area. The key insight underlying Magemite is, magnetic sensor integrated in smart devices can sense nearby magnetic field strength. Using a permanent magnet, users could “write” in ADS to communicate with matched devices. Different from previous magnetic-sensing schemes that recognize only coarse-grained gestures, Magemite can recognize user's fine-grained input like characters. However, individual's diverse writing patterns affect the recognition accuracy. To address this challenge, we preprocess the input trajectories and abstract different features of trajectories to uniquely identify user's input, then use these feature vectors to train several pattern recognition models for character recognition. We evaluate Magemite in various scenarios, and experimental results show Magemite can achieve average recognition accuracy over 85%. Yuyang Ke, Yan Xiong 0001, Yiqing Hu, Xudong Gong, Wenchao Huang 0001 |
WOWMOM | 3 |