EDBT 2026 Demo / reviewers in the wild / expert
Hui Su
dblp:89/1838
· DBLP profile ↗
95ranked-venue papers
21as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 9 first-author · 6 since 2021Computer networks · 9 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Forecasting of Geostationary Infrared Brightness Temperature Sequences: A Benchmark and a Lightweight ModelabstractForecasting geostationary infrared brightness temperature sequences from historical observations is a significant and challenging task. By analyzing these predictions, cloud evolution, convective activity, and atmospheric radiative states can be revealed in advance, offering high potential value in domains such as weather nowcasting, energy management, and disaster monitoring. Recently, artificial intelligence techniques have provided valuable insights into this task. However, as a nascent research area, the lack of a standardized, high-quality benchmark has significantly impeded progress. Moreover, training existing deep learning models for this task remains computationally expensive due to the complexity of their network architectures and modeling mechanisms. To address these challenges, we introduce a new benchmark, FY4ABT, and propose a lightweight prediction model, WavePredNet. Specifically, FY4ABT comprises three sub-datasets designed to respectively evaluate prediction performance under short-term, medium-term, and long-term scenarios. Meanwhile, WavePredNet effectively captures multi-scale dynamics, including both low- and high-frequency components with low computational costs while delivering exceptional performance. Kuai Dai, Hui Su, Xutao Li 0001, Chengxing Zhai |
AAAI | 2 |
| 2026 | SatSolarCast: A Flexible Framework for Multimodal Solar Irradiance Forecasting via Memory-Alignment LearningabstractSolar irradiance forecast aims to accurately estimate future solar irradiance based on historical data, playing a vital role in energy production and grid management. While ground-based station measurements provide local accuracy, geostationary satellites offer much broader environmental contexts, such as cloud coverage, which serves as a key factor for accurate forecasting. However, effectively integrating these multimodal observations remains a challenge, with existing methods suffering from inflexibility and high computational costs. To address this problem, we propose SatSolarCast, a flexible and efficient multimodal framework that introduces a memory alignment learning mechanism to integrate geostationary satellite data and historical irradiance observations. By preserving and recalling long-term spatiotemporal patterns from a specialized satellite memory bank, SatSolarCast enables effective guidance for both short- and long-term prediction. Additionally, SatSolarCast offers plug-and-play compatibility and can be incorporated into various forecasting architectures. Extensive experiments across four ground stations demonstrate that SatSolarCast substantially improves forecasting performance compared to prior methods with much lower computational costs. Kuai Dai, Hui Su, Chengxing Zhai, Huiwei Lin, Mingliang Bai |
AAAI | 2 |
| 2026 | Harnessing the Power of Reinforcement Learning for Language-Model-Based Information Retriever via Query-Document Co-Augmentation
Jingming Liu, Yao-Xiang Ding 0001, Hui Su, Kun Zhou 0001 |
PAKDD (3) | 5 |
| 2026 | Resource allocation for time-varying STAR-RIS aided NOMA systems with dynamic SIC decoding
Haining Liu, Xiaomeng Deng, Hui Su, Jie Jia 0001 |
Comput. Networks | 3 |
| 2026 | Dynamic reliable SFC orchestration for SDN-NFV enabled networks
Hui Su, Jie Jia 0001, Jian Chen 0008, Xingwei Wang 0001 |
Comput. Networks | 1 |
| 2026 | A hierarchical 5G-TSN integrated architecture with load-balancing oriented resource allocation
Hui Su, Jie Jia 0001, Jian Chen 0008, Xingwei Wang |
J. Netw. Comput. Appl. | 1 |
| 2025 | Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal InconsistencyabstractJiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang, Zheng Chu, Hui Su, Jinlan Fu, Ming Liu, See-Kiong Ng, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiafeng Liang, Shixin Jiang, Ning Wang 0020, Hui Su, Jinlan Fu, Ming Liu 0004, See-Kiong Ng, Bing Qin 0001 |
ACL (1) | 6 |
| 2025 | Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best PracticesabstractMultimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM. Junyan Lin, Yingqi Fan, Hui Su, Jinlan Fu |
CVPR | 6 |
| 2025 | Multimodal Language Models See Better When They Look ShallowerabstractMultimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT).This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis.While prior studies suggest that different ViT layers capture different types of information-shallower layers focusing on fine visual details and deeper layers aligning more closely with textual semantics, the impact of this variation on MLLM performance remains underexplored.We present the first comprehensive study of visual layer selection for MLLMs, analyzing representation similarity across ViT layers to establish shallow, middle, and deep layer groupings.Through extensive evaluation of MLLMs (1.4B-7B parameters) across 10 benchmarks encompassing 60+ tasks, we find that while deep layers excel in semantic-rich tasks like OCR, shallow and middle layers significantly outperform them on fine-grained visual tasks including counting, positioning, and object localization.Building on these insights, we propose a lightweight feature fusion method that strategically incorporates shallower layers, achieving consistent improvements over both single-layer and specialized fusion baselines.Our work offers the first principled study of visual layer selection in MLLMs, showing that MLLMs can often see better when they look shallower. Junyan Lin, Xinghao Chen 0009, Jianfeng Dong, Xin Jin 0014, Hui Su, Jinlan Fu, Xiaoyu Shen 0001 |
EMNLP | 7 |
| 2025 | VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) have achieved strong performance across vision-language tasks, but suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens.Though efforts have been made to prune tokens in MLLMs, they lack a fundamental understanding of how MLLMs process and fuse multimodal information.Through systematic analysis, we uncover a three-stage cross-modal interaction process: (1) Shallow layers recognize task intent, with visual tokens acting as passive attention sinks; (2) Cross-modal fusion occurs abruptly in middle layers, driven by a few critical visual tokens; (3) Deep layers discard vision tokens, focusing solely on linguistic refinement.Based on these findings, we propose VisiPruner, a training-free pruning framework that reduces up to 99% of visionrelated attention computations and 53.9% of FLOPs on LLaVA-v1.5 7B.It significantly outperforms existing token pruning methods and generalizes across diverse MLLMs.Beyond pruning, our insights further provide actionable guidelines for training efficient MLLMs by aligning model architecture with its intrinsic layer-wise processing dynamics. Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang 0185, Xiaoyu Shen 0001 |
EMNLP | 5 |
| 2025 | PSCA: A FPGA-based Protein Structure Comparison Accelerator with Symmetric Simplified Matrix
Hui Su, Xingyun Qi, Qiang Wang 0006, Puguang Liu, Haoyu Liao |
ICA3PP (6) | 1 |
| 2025 | SkipGPT: Each Token is One of a KindabstractLarge language models (LLMs) achieve remarkable performance across tasks but incur substantial computational costs due to their deep, multi-layered architectures. Layer pruning has emerged as a strategy to alleviate these inefficiencies, but conventional static pruning methods overlook two critical dynamics inherent to LLM inference: (1) *horizontal dynamics*, where token-level heterogeneity demands context-aware pruning decisions, and (2) *vertical dynamics*, where the distinct functional roles of MLP and self-attention layers necessitate component-specific pruning policies. We introduce **SkipGPT**, a dynamic layer pruning framework designed to optimize computational resource allocation through two core innovations: (1) global token-aware routing to prioritize critical tokens and (2) decoupled pruning policies for MLP and self-attention components. To mitigate training instability, we propose a two-stage optimization paradigm: first, a disentangled training phase that learns routing strategies via soft parameterization to avoid premature pruning decisions, followed by parameter-efficient LoRA fine-tuning to restore performance impacted by layer removal. Extensive experiments demonstrate that SkipGPT reduces over 40% model parameters while matching or exceeding the performance of the original dense model across benchmarks. By harmonizing dynamic efficiency with preserved expressivity, SkipGPT advances the practical deployment of scalable, resource-aware LLMs. Our code is publicly available at: https://github.com/EIT-NLP/SkipGPT. Anhao Zhao, Fanghua Ye 0001, Yingqi Fan, Junlong Tong, Zhiwei Fei, Hui Su, Xiaoyu Shen 0001 |
ICML | 7 |
| 2025 | 3D-MGW: A Memory-Efficient Grouped Watermark for Multi-Object 3D Gaussian SplattingabstractMulti-object 3D Gaussian Splatting (3DGS) technology aims to efficiently synthesize complex 3D scenes from images while allowing users to manipulate objects through textual prompts. Training multi-object 3DGS models requires substantial computational resources, making it necessary to protect the generated 3D objects from unauthorized reproduction, modification, and distribution. Existing watermarking solutions suffer from high memory consumption and require additional time overhead. Moreover, they cannot precisely localize watermarks to specific objects, making it difficult to trace individual contributions when multiple creators collaborate on a same multi-object scene. To address these challenges, we propose 3D-MGW, a novel memory-efficient grouped watermark for multi-object 3DGS. Within 3D-MGW, background scenes are reconstructed from images, while diffusion models are incorporated to guide highquality 3DGS synthesis from prompts. To eliminate additional training overhead, watermark embedding is integrated within the 3DGS training process rather than implementing it separately. Subsequently, a grouped Gaussian strategy is introduced to enable granular, high-capacity multi-object watermark. Additionally, a Gaussian compression module is proposed to eliminate redundant primitives, reducing the storage footprint of Gaussian models. Through comprehensive experiments, our 3D-MGW demonstrates exceptional performance, achieving 95% watermark extraction accuracy under 64-bit watermarks while reducing storage utilization by up to 61%, highlighting its substantial potential for multi-object 3DGS applications. Hui Su, Gaolei Li, Wenkai Huang 0003, Xiaoyu Yi 0003, Jianhua Li 0001 |
ICPADS | 1 |
| 2025 | Multi-Authority CP-ABE Scheme With Cryptographic Reverse Firewalls for Internet of VehiclesabstractInternet of vehicles, featured with widely distributed vehicle nodes and limited computing power, usually have high performance requirements. Because of this feature, efficient and reliable access control has raised a challenge in Internet of vehicles. Ciphertext-policy attribute-based encryption (CP-ABE) could be denoted as an efficient solution for this problem. However, directly applying traditional single-authority CP-ABE schemes may result in single-point performance bottleneck. Besides, the secrets of the whole system may be leaked if any node is attacked. To solve these challenging tasks, we proposed MA-CP-ABE-CRF, a multi-authority CP-ABE scheme with cryptographic reverse firewalls. The system is designed to grant vehicles fine-grained access control by encrypting data under vehicle attributes. Besides, load balancing of authorization in distributed systems is achieved based on the characteristic of multi-authority. Meanwhile, specific nodes are equipped with cryptographic reverse firewalls (CRFs) to prevent information leakage. As the first scheme with the above features for Internet of vehicles, the system achieves adaptive CPA-security and ASA-security. Through rigorous theoretical analysis and experimental comparison, MA-CP-ABE-CRF is proved to be highly efficient and practical. Hu Xiong, Hui Su, Kuo-Hui Yeh |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual ErrorsabstractAccurate text summarization is one of the most common and important tasks performed by Large Language Models, where the costs of human review for an entire document may be high, but the costs of errors in summarization may be even greater.We propose Detecting Errors through Ensembling Prompts (DEEP) -an end-to-end large language model framework for detecting factual errors in text summarization.Our framework uses a diverse set of LLM prompts to identify factual inconsistencies, treating their outputs as binary features, which are then fed into ensembling models.We then calibrate the ensembled models to produce empirically accurate probabilities that a text is factually consistent or free of hallucination.We demonstrate that prior models for detecting factual errors in summaries perform significantly worse without optimizing the thresholds on subsets of the evaluated dataset.Our framework achieves state-of-the-art (SOTA) balanced accuracy on the AggreFact-XSUM FTSOTA, To-fuEval Summary-Level, and HaluEval Summarization benchmarks in detecting factual errors within transformer-generated text summaries.It does so without any fine-tuning of the language model or reliance on thresholding techniques not available in practical settings.1 Alex Chandler, Devesh Surve, Hui Su |
EMNLP | 3 |
| 2024 | A High-performance Hardware Accelerator for Genome AlignmentabstractGenome alignment is a vital process in genome sequencing and bio-informatics research. It entails aligning short DNA sequence fragments (usually tens to hundreds of base pairs) with a reference genome sequence, uncovering crucial biological information and variations. However, due to the high computational complexity of current sequence matching algorithms and the rapid growth of genetic data, there exists a computational bottleneck in sequence alignment workflows. Therefore, scholars have turned to using hardware to accelerate this computation, with a focus on high-performance computing. While existing acceleration solutions have mostly concentrated on the classical algorithm, showing significant improvements, there has been limited work on accelerating alignment algorithms in popular bio-informatics software tools (enhanced version). In this paper, we proposed a hardware accelerator for the alignment tasks in the popular sequence alignment tool minimap2 using the KSW2 algorithm. Our approach utilizes an anti-diagonal processing element(PE) array for parallel computation, implements the BAND technology in hardware, and enables support for longer input sequences without sacrificing accuracy. Additionally, we integrated the traceback stage of the KSW2 algorithm in hardware, reducing communication data overhead. We attained a 15.32x acceleration compared to software optimized with Streaming SIMD Extensions (SSE) instruction sets and hyper-threading technology. Additionally, our design shows a 3.42x speedup compared to other high-performance hardware. Haoyu Liao, Hui Su, Xingyun Qi |
ISPA | 5 |
| 2023 | SASFormer: Transformers for Sparsely Annotated Semantic SegmentationabstractSemantic segmentation based on sparse annotation has advanced in recent years. It labels only part of each object in the image, leaving the remainder unlabeled. Most of the existing approaches are time-consuming and often necessitate a multi-stage training strategy. In this work, we propose a simple yet effective sparse annotated semantic segmentation framework based on segformer, dubbed SASFormer, that achieves remarkable performance. Specifically, the framework first generates hierarchical patch attention maps, which are then multiplied by the network predictions to produce correlated regions separated by valid labels. Besides, we also introduce the affinity loss to ensure consistency between the features of correlation results and network predictions. Extensive experiments showcase that our proposed approach is superior to existing methods and achieves cutting-edge performance. The source code is available at https://github.com/su-hui-zz/SASFormer. Hui Su, Yue Ye, Lechao Cheng, Mingli Song |
ICME | 1 |
| 2023 | ACETest: Automated Constraint Extraction for Testing Deep Learning OperatorsabstractDeep learning (DL) applications are prevalent nowadays as they can help with multiple tasks. DL libraries are essential for building DL applications. Furthermore, DL operators are the important building blocks of the DL libraries, that compute the multi-dimensional data (tensors). Therefore, bugs in DL operators can have great impacts. Testing is a practical approach for detecting bugs in DL operators. In order to test DL operators effectively, it is essential that the test cases pass the input validity check and are able to reach the core function logic of the operators. Hence, extracting the input validation constraints is required for generating high-quality test cases. Existing techniques rely on either human effort or documentation of DL library APIs to extract the constraints. They cannot extract complex constraints and the extracted constraints may differ from the actual code implementation. To address the challenge, we propose ACETest, a technique to automatically extract input validation constraints from the code to build valid yet diverse test cases which can effectively unveil bugs in the core function logic of DL operators. For this purpose, ACETest can automatically identify the input validation code in DL operators, extract the related constraints and generate test cases according to the constraints. The experimental results on popular DL libraries, TensorFlow and PyTorch, demonstrate that ACETest can extract constraints with higher quality than state-of-the-art (SOTA) techniques. Moreover, ACETest is capable of extracting 96.4% more constraints and detecting 1.95 to 55 times more bugs than SOTA techniques. In total, we have used ACETest to detect 108 previously unknown bugs on TensorFlow and PyTorch, with 87 of them confirmed by the developers. Lastly, five of the bugs were assigned with CVE IDs due to their security impacts. Yang Xiao 0011, Yuekang Li, Yeting Li, Dongsong Yu, Chendong Yu, Hui Su, Wei Huo 0005 |
ISSTA | 7 |
| 2023 | A survey of transformer-based multimodal pre-trained modals
Xue Han 0018, Junlan Feng, Chao Deng 0002, Hui Su, Lun Hu, Pengwei Hu 0001 |
Neurocomputing | 7 |
| 2023 | Analysis of ENF Signal Extraction From Videos Acquired by Rolling ShuttersabstractElectric network frequency (ENF) analysis is a promising forensic technique for authenticating multimedia recordings and detecting tampering. The validity of the ENF analysis heavily relies on the capability of extracting high-quality ENF signals from multimedia recordings. This paper analyzes and compares two representative methods for extracting ENF signals from visual signals acquired by cameras using the rolling-shutter mechanism. The first method proposed in prior work,direct concatenation, ignores the idle period of each frame. The second method proposed in this paper,periodic zeroing-out, inserts zeros to missing sample points instead of ignoring the idle period. Our theoretical analyses of using multirate signal processing reveal and experiments confirm that while the first method can extract ENF signals without knowing the exact value of camera read-out time, there exists some mild distortion to extracted ENF signals. In contrast, the second method taking the read-out time as the additional input is capable of extracting distortion-free ENF signals, and its frequency component of the highest strength is always located at the nominal frequency. Additionally, we examine aliased DC and negative ENF components caused by the two methods and show that their impact on the accuracy of frequency estimation is minimum. This paper facilitates the fundamental understanding of extracting ENF signals from videos. The research findings imply that the periodic zeroing-out method offers more accurate frequency estimates, but the performance improvement is not significant. Jisoo Choi, Chau-Wai Wong, Hui Su, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Parallel Training of Pre-Trained Models via Chunk-Based Dynamic Memory ManagementabstractThe pre-trained model (PTM) is revolutionizing Artificial Intelligence (AI) technology. However, the hardware requirement of PTM training is prohibitively high, making it a game for a small proportion of people. Therefore, we proposed PatrickStar system to lower the hardware requirements of PTMs and make them accessible to everyone. PatrickStar uses the CPU-GPU heterogeneous memory space to store the model data. Different from existing works, we organize the model data in memory chunks and dynamically distribute them in the heterogeneous memory. Guided by the runtime memory statistics collected in a warm-up iteration, chunks are orchestrated efficiently in heterogeneous memory and generate lower CPU-GPU data transmission volume and higher bandwidth utilization. Symbiosis with the Zero Redundancy Optimizer, PatrickStar scales to multiple GPUs on multiple nodes. The system can train tasks on bigger models and larger batch sizes, which cannot be accomplished by existing works. Experimental results show that PatrickStar extends model scales 2.27 and 2.5 times of DeepSpeed, and exhibits significantly higher execution speed. PatricStar also successfully runs the 175B GPT3 training task on a 32 GPU cluster. Our code is available athttps://github.com/Tencent/PatrickStar. Jiarui Fang, Zilin Zhu, Shenggui Li, Hui Su, Yang Yu 0038, Jie Zhou 0016, Yang You 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingabstractLarge-scale pretrained language models have achieved SOTA results on NLP tasks.However, they have been shown vulnerable to adversarial attacks especially for logographic languages like Chinese.In this work, we propose ROCBERT: a pretrained Chinese Bert that is robust to various forms of adversarial attacks like word perturbation, synonyms, typos, etc.It is pretrained with the contrastive learning objective which maximizes the label consistency under different synthesized adversarial examples.The model takes as input multimodal information including the semantic, phonetic and visual features.We show all these features are important to the model robustness since the attack can be performed in all the three forms.Across 5 Chinese NLU tasks, ROCBERT outperforms strong baselines under three blackbox adversarial algorithms without sacrificing the performance on clean testset.It also performs the best in the toxic content detection task under human-made attacks. * Equal contribution. Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Tuo Ji, Jiarui Fang, Jie Zhou 0016 |
ACL (1) | 1 |
| 2022 | Re-Attention Transformer for Weakly Supervised Object Localization
Hui Su, Yue Ye, Mingli Song, Lechao Cheng |
BMVC | 1 |
| 2022 | Dynamic Targeting for Improved Tracking of Storm FeaturesabstractDynamic Targeting (DT) will enable future Earth Observing instruments to intelligently reconfigure and point instruments to dramatically enhance science return. In this work we present a realistic simulation study of DT for tracking of storm features. To this end we have developed several algorithms from Operations Research and Artificial Intelli-gence/heuristic search. We benchmark these algorithms and show that DT is a powerful tool with the potential to significantly improve science yield. Alberto Candela, Jason Swope, Steve A. Chien, Hui Su, Peyman Tavallali |
IGARSS | 4 |
| 2022 | Congestion-Aware Modeling and Analysis of Sponsored Data Plan from End User PerspectiveabstractThe past decade has witnessed the rapid expansion of demands for mobile traffic, while the traditional mobile traffic pricing schemes cannot accommodate such demands. Sponsored data plan (SDP), which can increase the revenue of all stakeholders in the market through transferring some of the revenue from content providers (CPs) to end users (EUs), is more suitable. However, existing studies have focused more on Internet service providers (ISPs) and CPs, ignoring the influence of EUs (e.g., the inherent attribute differences of EUs and the interaction among EUs) on the market under SDP. Regarding the difficulty of modeling the abstract property about interaction among EUs, we utilize network congestion as the medium and construct the congestion-aware SDP model based on Stackelberg game. The newly proposed model can not only analyze how network congestion affects SDP mechanism, but also elucidate the impact of interactions among EUs. More specifically, through theoretical analysis, we prove that there is a unique dynamic equilibrium in the interaction among EUs (i.e., the traffic consumption of different EUs). By taking into account network congestion, the newly proposed model also more accurately and realistically describes the optimal strategies and computation methods of all stakeholders in the market. Moreover, simulation experiments demonstrate that the positive effect brought by SDP is not as obvious as before, and EUs influence each other instead of being independent of each other. Overall, this paper emphasizes the non-negligible influence of EUs and promotes a deeper understanding of SDP mechanism, which can guide the relevant stakeholders to optimize their own decision-making details. Yi Zhao 0011, Qi Tan 0003, Xiaohua Xu 0002, Hui Su, Dan Wang 0002, Ke Xu 0002 |
IWQoS | 4 |
| 2022 | Temporal Multimodal Multivariate LearningabstractWe introduce temporal multimodal multivariate learning, a new family of decision making models that can indirectly learn and transfer online information from simultaneous observations of a probability distribution with more than one peak or more than one outcome variable from one time stage to another. We approximate the posterior by sequentially removing additional uncertainties across different variables and time, based on data-physics driven correlation, to address a broader class of challenging time-dependent decision-making problems under uncertainty. Extensive experiments on real-world datasets ( i.e., urban traffic data and hurricane ensemble forecasting data) demonstrate the superior performance of the proposed targeted decision-making over the state-of-the-art baseline prediction methods across various settings. Hyoshin Park, Justice Darko, Niharika Deshpande, Venktesh Pandey, Hui Su, Masahiro Ono, Dedrick Barkely, Larkin Folsom, Derek J. Posselt, Steve A. Chien |
KDD | 5 |
| 2021 | Neural Data-to-Text Generation with LM-based Text AugmentationabstractFor many new application domains for datato-text generation, the main obstacle in training neural models consists of a lack of training data.While usually large numbers of instances are available on the data side, often only very few text samples are available.To address this problem, we here propose a novel fewshot approach for this setting.Our approach automatically augments the data available for training by (i) generating new text samples based on replacing specific values by alternative ones from the same category, (ii) generating new text samples based on GPT-2, and (iii) proposing an automatic method for pairing the new text samples with data samples.As the text augmentation can introduce noise to the training data, we use cycle consistency as an objective, in order to make sure that a given data sample can be correctly reconstructed after having been formulated as text (and that text samples can be reconstructed from data).On both the E2E and WebNLG benchmarks, we show that this weakly supervised training paradigm is able to outperform fully supervised seq2seq models with less than 10% annotations.By utilizing all annotated data, our model can boost the performance of a standard seq2seq model by over 5 BLEU points, establishing a new state-of-the-art on both datasets. * Work done prior to joining Amazon.The Blue Spice is a restaurant that serves English cuisine. Ernie Chang, Xiaoyu Shen 0001, Vera Demberg, Hui Su |
EACL | 5 |
| 2021 | Complementary Evidence Identification in Open-Domain Question AnsweringabstractThis paper proposes a new problem of complementary evidence identification for opendomain question answering (QA).The problem aims to efficiently find a small set of passages that covers full evidence from multiple aspects as to answer a complex question.To this end, we proposes a method that learns vector representations of passages and models the sufficiency and diversity within the selected set, in addition to the relevance between the question and passages.Our experiments demonstrate that our method considers the dependence within the supporting evidence and significantly improves the accuracy of complementary evidence selection in QA domain. Xiangyang Mou, Mo Yu, Shiyu Chang, Hui Su |
EACL | 6 |
| 2021 | Improved Chroma from Luma Prediction in AV1 Based On Virtual Chroma Block GenerationabstractChroma from Luma (CfL) prediction is an efficient coding tool in AV1 which builds chroma prediction by implementing a linear model on luma pixels. To avoid transmission, the off-set factor in the linear model is set to the average of neighboring chroma pixels. An improved CfL algorithm is proposed to derive the offset factor based on a virtual chroma block. Such a block is constructed by using the chroma of the matched pixel which is determined according to the luma difference of neighboring regions and the co-located luma. Compared with libaom, the proposed CfL algorithm provides 0.31% and 0.15% weighted PSNR BD-rate saving under AI and RA configuration, respectively. Experimental results show that over 1.00% and 0.80% BD-rate saving can be achieved for chroma components under AI and RA configuration. With the proposed algorithm, the percentage of pixels with CfL as the optimal coding mode is increased. Junyan Huo, Menglin Zhang, Wenhan Qiao, Fuzheng Yang 0001, Hui Su, Debargha Mukherjee |
ICME | 5 |
| 2021 | A Technical Overview of AV1abstractThe AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than a 30% reduction in bit rate compared to its predecessor VP9 for the same decoded video quality. This article provides a technical overview of the AV1 codec design that enables the compression performance gains with considerations for hardware feasibility. Jingning Han, Bohan Li 0006, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, Yue Chen 0040, Yunqing Wang, Paul Wilkins, Yaowu Xu, Jim Bankoski |
Proc. IEEE | 7 |
| 2021 | Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive StudyabstractAbstract Recent advancements in open-domain question answering (ODQA), that is, finding answers from large open-domain corpus like Wikipedia, have led to human-level performance on many datasets. However, progress in QA over book stories (Book QA) lags despite its similar task formulation to ODQA. This work provides a comprehensive and quantitative analysis about the difficulty of Book QA: (1) We benchmark the research on the NarrativeQA dataset with extensive experiments with cutting-edge ODQA techniques. This quantifies the challenges Book QA poses, as well as advances the published state-of-the-art with a ∼7% absolute improvement on ROUGE-L. (2) We further analyze the detailed challenges in Book QA through human studies.1 Our findings indicate that the event-centric questions dominate this task, which exemplifies the inability of existing QA models to handle event-oriented scenarios. Xiangyang Mou, Chenghao Yang 0001, Mo Yu, Bingsheng Yao, Saloni Potdar, Hui Su |
Trans. Assoc. Comput. Linguistics | 7 |
| 2020 | Neural Data-to-Text Generation via Jointly Learning the Segmentation and CorrespondenceabstractThe neural attention model has achieved great success in data-to-text generation tasks. Though usually excelling at producing fluent text, it suffers from the problem of information missing, repetition and "hallucination". Due to the black-box nature of the neural attention architecture, avoiding these problems in a systematic way is non-trivial. To address this concern, we propose to explicitly segment target text into fragment units and align them with their data correspondences. The segmentation and correspondence are jointly learned as latent variables without any human annotations. We further impose a soft statistical constraint to regularize the segmental granularity. The resulting architecture maintains the same expressive power as neural attention models, while being able to generate fully interpretable outputs with several times less computational cost. On both E2E and WebNLG benchmarks, we show the proposed model consistently outperforms its neural attention counterparts. Xiaoyu Shen 0001, Ernie Chang, Hui Su, Cheng Niu, Dietrich Klakow |
ACL | 3 |
| 2020 | Diversifying Dialogue Generation with Non-Conversational TextabstractNeural network-based sequence-to-sequence (seq2seq) models strongly suffer from the lowdiversity problem when it comes to opendomain dialogue generation.As bland and generic utterances usually dominate the frequency distribution in our daily chitchat, avoiding them to generate more interesting responses requires complex data filtering, sampling techniques or modifying the training objective.In this paper, we propose a new perspective to diversify dialogue generation by leveraging non-conversational text.Compared with bilateral conversations, nonconversational text are easier to obtain, more diverse and cover a much broader range of topics.We collect a large-scale nonconversational corpus from multi sources including forum comments, idioms and book snippets.We further present a training paradigm to effectively incorporate these text via iterative back translation.The resulting model is tested on two conversational datasets and is shown to produce significantly more diverse responses without sacrificing the relevance with context. Hui Su, Xiaoyu Shen 0001, Sanqiang Zhao, Xiao Zhou 0004, Pengwei Hu 0001, Randy Zhong, Cheng Niu, Jie Zhou 0016 |
ACL | 1 |
| 2020 | Bayesian Adversarial Human Motion SynthesisabstractWe propose a generative probabilistic model for human motion synthesis. Our model has a hierarchy of three layers. At the bottom layer, we utilize Hidden semi-Markov Model (HSMM), which explicitly models the spatial pose, temporal transition and speed variations in motion sequences. At the middle layer, HSMM parameters are treated as random variables which are allowed to vary across data instances in order to capture large intra- and inter-class variations. At the top layer, hyperparameters define the prior distributions of parameters, preventing the model from overfitting. By explicitly capturing the distribution of the data and parameters, our model has a more compact parameterization compared to GAN-based generative models. We formulate the data synthesis as an adversarial Bayesian inference problem, in which the distributions of generator and discriminator parameters are obtained for data synthesis. We evaluate our method through a variety of metrics, where we show advantage than other competing methods with better fidelity and diversity. We further evaluate the synthesis quality as a data augmentation method for recognition task. Finally, we demonstrate the benefit of our fully probabilistic approach in data restoration task. Rui Zhao 0015, Hui Su |
CVPR | 2 |
| 2020 | MovieChats: Chat like Humans in a Closed DomainabstractBeing able to perform in-depth chat with humans in a closed domain is a precondition before an open-domain chatbot can ever be claimed.In this work, we take a close look at the movie domain and present a large-scale high-quality corpus with fine-grained annotations in hope of pushing the limit of moviedomain chatbots.We propose a unified, readily scalable neural approach which reconciles all subtasks like intent prediction and knowledge retrieval.The model is first pretrained on the huge general-domain data, then finetuned on our corpus.We show this simple neural approach trained on high-quality data is able to outperform commercial systems replying on complex rules.On both the static and interactive tests, we find responses generated by our system exhibits remarkably good engagement and sensibleness close to human-written ones.We further analyze the limits of our work and point out potential directions for future work 1 . Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Ernie Chang, Cheng Niu, Jie Zhou 0016 |
EMNLP (1) | 1 |
| 2020 | Machine Learning Based Symbol Probability Distribution Prediction For Entropy Coding In Av1abstractEntropy coding is a lossless data compression technique that is widely applied in video codecs to encode syntax elements into bitstreams. Efficient entropy coding requires accurate prediction of the probability distribution of the encoded symbols. In AV1, multi-symbol arithmetic coding is adopted. The symbol probability is derived with handcrafted context models and lookup tables that store the predicted probabilities corresponding to different entropy contexts. The lookup table based scheme has some fundamental deficiencies. The entropy context features have to be discrete so that they can be used to index the lookup tables. To reduce the size of the lookup table, the number of contexts cannot be very large. Moreover, the probability distributions stored in the lookup tables are maintained separately without taking their correlations into consideration. In this paper, we propose a machine learning based scheme that achieves more accurate symbol probability prediction for entropy coding. The proposed approach is implemented in AV1 for the entropy coding of intra prediction modes. Experimental results demonstrate that it can improve the efficiency of entropy coding significantly. Mingliang Chen 0001, Hui Su, Sai Deng, Yaowu Xu |
ICIP | 2 |
| 2020 | BlueMemo: Depression Analysis through Twitter PostsabstractThe use of social media runs through our lives, and users' emotions are also affected by it. Previous studies have reported social organizations and psychologists using social media to find depressed patients. However, due to the variety of content published by users, it isn't effortless for the system to consider the text, image, and even the hidden information behind the image. To address this problem, we proposed a new system for social media screening of depressed patients named BlueMemo. We collected real-time posts from Twitter. Based on the posts, learned text features, image features, and visual attributes were extracted as three modalities and were fed into a multi-modal fusion and classification model to implement our system. The proposed BlueMemo has the power to help physicians and clinicians quickly and accurately identify users at potential risk for depression. Pengwei Hu 0001, Chenhao Lin, Hui Su, Shaochun Li, Xue Han 0018, Jing Mei |
IJCAI | 3 |
| 2020 | A Collaborative, Immersive Language Learning Environment Using Augmented Panoramic ImageryabstractAugmenting immersive technologies with AI for foreign language learning is a relatively unexplored, multidisciplinary, and complex research paradigm. The Cognitive and Immersive Room at Rensselaer Polytechnic Institute facilitates cultural and foreign language learning beyond a traditional classroom setting. Exploration of virtual environments and real-world renderings occur at human-scale by groups of students and teachers simultaneously, without the need for head-mounted displays. Students travel to true-to-life Panoramic Scenes to learn authentic cultural knowledge and practice vocabulary within the same surroundings they would use such knowledge while traveling abroad. Multimodal input allows students to engage with the system through gesture, voice, and spatial positioning, creating a dynamic language learning experience. This work underwent initial assessment in a six-week summer course, AI-Assisted Immersive Chinese, and preliminary qualitative results find a majority of surveyed students indicate the system is useful, engaging, and fun. Samuel Chabot, Jaimie Drozdal, Matthew Peveler, Yalun Zhou, Hui Su, Jonas Braasch |
iLRN | 5 |
| 2020 | Trust in AutoML: exploring information needs for establishing trust in automated machine learning systemsabstractWe explore trust in a relatively new area of data science: Automated Machine Learning (AutoML). In AutoML, AI methods are used to generate and optimize machine learning models by automatically engineering features, selecting models, and optimizing hyperparameters. In this paper, we seek to understand what kinds of information influence data scientists' trust in the models produced by AutoML? We operationalize trust as a willingness to deploy a model produced using automated methods. We report results from three studies - qualitative interviews, a controlled experiment, and a card-sorting task - to understand the information needs of data scientists for establishing trust in AutoML systems. We find that including transparency features in an AutoML tool increased user trust and understandability in the tool; and out of all proposed features, model performance metrics and visualizations are the most important information to data scientists when establishing their trust with an AutoML tool. Jaimie Drozdal, Justin D. Weisz, Dakuo Wang, Gaurav Dass, Bingsheng Yao, Changruo Zhao, Michael J. Muller, Lin Ju, Hui Su |
IUI | 9 |
| 2020 | SenseMood: Depression Detection on Social MediaabstractMore than 300 million people have been affected by depression all over the world. Due to the medical equipment and knowledge limitations, most of them are not diagnosed at the early stages. Recent work attempts to use social media to detect depression since the patterns of opinions and thoughts expression of the posted text and images, can reflect users' mental state to some extent. In this work, we design a system dubbed SenseMood to demonstrate that the users with depression can be efficiently detected and analyzed by using proposed system. A deep visual-textual multimodal learning approach has been proposed to reveal the psychological state of the users on social networks. The posted images and tweets data from users with/without depression on Twitter have been collected and used for depression detection. CNN-based classifier and Bert are applied to extract the deep features from the pictures and text posted by users respectively. Then visual and textual features are combined to reflect the emotional expression of users. Finally our system classifies the users with depression and normal users through a neural network and the analysis report is generated automatically. Chenhao Lin, Pengwei Hu 0001, Hui Su, Shaochun Li, Jing Mei, Jie Zhou 0016, Henry Leung 0001 |
ICMR | 3 |
| 2020 | Incentive mechanisms for mobile data offloading through operator-owned WiFi access points
Yi Zhao 0011, Ke Xu 0002, Yifeng Zhong, Xiang-Yang Li 0001, Ning Wang 0001, Hui Su, Meng Shen 0001 |
Comput. Networks | 6 |
| 2020 | Understand Love of Variety in Wireless Data Market Under Sponsored Data PlansabstractSponsored Data Plan (SDP) is an emerging pricing model for the wireless data market where the Content Provider (CP) can sponsor the data usage for specific content on behalf of the users. This strategy sheds new light on the data pricing model and receives significant attention from the Internet Service Provider (ISP). However, the existing SDP studies consider traffic price (e.g., sponsorship) as the only factor that affects user decision. The impact of other classic market features, such as the demand for a variety of contents (i.e., love of variety), remains largely unclear. In this paper, we develop a new model to understand the love of variety in the wireless data market under SDPs. Our model has demonstrated that, such variety is important to understand the complex gaming between ISPs, CPs, and users in both short-run and long-run markets. For example, the analysis indicates that the advantage of CPs with higher revenue will be significantly reduced when users have a greater love of variety. Moreover, to help the ISP better adopt the proposed model in the real market, we also develop a practical method to calibrate the related parameters, which can also be applied to quantity the love of variety. Yi Zhao 0011, Hui Su, Liang Zhang 0042, Rui Zhang 0017, Dan Wang 0002, Ke Xu 0002 |
IEEE J. Sel. Areas Commun. | 3 |
| 2020 | Joint deep semantic embedding and metric learning for person re-identification
Yan-shuo Chang, Hui Su, Ni Gao, Xin-An Yang |
Pattern Recognit. Lett. | 5 |
| 2019 | The Rensselaer Mandarin Project - A Cognitive and Immersive Language Learning EnvironmentabstractThe Rensselaer Mandarin Project enables a group of foreign language students to improve functional understanding, pronunciation and vocabulary in Mandarin Chinese through authentic speaking situations in a virtual visit to China. Students use speech, gestures, and combinations thereof to navigate an immersive, mixed reality, stylized realism game experience through interaction with AI agents, immersive technologies, and game mechanics. The environment was developed in a black box theater equipped with a human-scale 360◦ panoramic screen (140h, 200r), arrays of markerless motion tracking sensors, and speakers for spatial audio. Rahul R. Divekar, Jaimie Drozdal, Lilit Balagyozyan, Shuyue Zheng, Ziyi Song, Huang Zou, Jeramey Tyler, Xiangyang Mou, Rui Zhao 0015, Helen Zhou, Jianling Yue, Jeffrey O. Kephart, Hui Su |
AAAI | 14 |
| 2019 | Improving Multi-turn Dialogue Modelling with Utterance ReWriterabstractRecent research has achieved impressive results in single-turn dialogue modelling. In the multi-turn setting, however, current models are still far from satisfactory. One major challenge is the frequently occurred coreference and information omission in our daily conversation, making it hard for machines to understand the real intention. In this paper, we propose rewriting the human utterance as a pre-process to help multi-turn dialgoue modelling. Each utterance is first rewritten to recover all coreferred and omitted information. The next processing steps are then performed based on the rewritten utterance. To properly train the utterance rewriter, we collect a new dataset with human annotations and introduce a Transformer-based utterance rewriting architecture using the pointer network. We show the proposed architecture achieves remarkably good performance on the utterance rewriting task. The trained utterance rewriter can be easily integrated into online chatbots and brings general improvement over different domains. Hui Su, Xiaoyu Shen 0001, Rongzhi Zhang, Fei Sun 0001, Pengwei Hu 0001, Cheng Niu, Jie Zhou 0016 |
ACL (1) | 1 |
| 2019 | Neuro-Inspired Eye Tracking With Eye Movement DynamicsabstractGeneralizing eye tracking to new subjects/environments remains challenging for existing appearance-based methods. To address this issue, we propose to leverage on eye movement dynamics inspired by neurological studies. Studies show that there exist several common eye movement types, independent of viewing contents and subjects, such as fixation, saccade, and smooth pursuits. Incorporating generic eye movement dynamics can therefore improve the generalization capabilities. In particular, we propose a novel Dynamic Gaze Transition Network (DGTN) to capture the underlying eye movement dynamics and serve as the topdown gaze prior. Combined with the bottom-up gaze measurements from the deep convolutional neural network, our method achieves better performance for both within-dataset and cross-dataset evaluations compared to state-of-the-art. In addition, a new DynamicGaze dataset is also constructed to study eye movement dynamics and eye gaze estimation. Kang Wang 0002, Hui Su |
CVPR | 2 |
| 2019 | Generalizing Eye Tracking With Bayesian Adversarial LearningabstractExisting appearance-based gaze estimation approaches with CNN have poor generalization performance. By systematically studying this issue, we identify three major factors: 1) appearance variations; 2) head pose variations and 3) over-fitting issue with point estimation. To improve the generalization performance, we propose to incorporate adversarial learning and Bayesian inference into a unified framework. In particular, we first add an adversarial component into traditional CNN-based gaze estimator so that we can learn features that are gaze-responsive but can generalize to appearance and pose variations. Next, we extend the point-estimation based deterministic model to a Bayesian framework so that gaze estimation can be performed using all parameters instead of only one set of parameters. Besides improved performance on several benchmark datasets, the proposed method also enables online adaptation of the model to new subjects/environments, demonstrating the potential usage for practical real-time eye tracking applications. Kang Wang 0002, Rui Zhao 0015, Hui Su |
CVPR | 3 |
| 2019 | Bayesian Hierarchical Dynamic Model for Human Action RecognitionabstractHuman action recognition remains as a challenging task partially due to the presence of large variations in the execution of action. To address this issue, we propose a probabilistic model called Hierarchical Dynamic Model (HDM). Leveraging on Bayesian framework, the model parameters are allowed to vary across different sequences of data, which increase the capacity of the model to adapt to intra-class variations on both spatial and temporal extent of actions. Meanwhile, the generative learning process allows the model to preserve the distinctive dynamic pattern for each action class. Through Bayesian inference, we are able to quantify the uncertainty of the classification, providing insight during the decision process. Compared to state-of-the-art methods, our method not only achieves competitive recognition performance within individual dataset but also shows better generalization capability across different datasets. Experiments conducted on data with missing values also show the robustness of the proposed method. Rui Zhao 0015, Wanru Xu, Hui Su |
CVPR | 3 |
| 2019 | Select and Attend: Towards Controllable Content Selection in Text GenerationabstractXiaoyu Shen, Jun Suzuki, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaoyu Shen 0001, Jun Suzuki 0001, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Improving Latent Alignment in Text Summarization by Generalizing the Pointer GeneratorabstractXiaoyu Shen, Yang Zhao, Hui Su, Dietrich Klakow. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaoyu Shen 0001, Hui Su, Dietrich Klakow |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Face Alignment With Kernel Density Deep Neural NetworkabstractDeep neural networks achieve good performance in many computer vision problems such as face alignment. However, when the testing image is challenging due to low resolution, occlusion or adversarial attacks, the accuracy of a deep neural network suffers greatly. Therefore, it is important to quantify the uncertainty in its predictions. A probabilistic neural network with Gaussian distribution over the target is typically used to quantify uncertainty for regression problems. However, in real-world problems especially computer vision tasks, the Gaussian assumption is too strong. To model more general distributions, such as multi-modal or asymmetric distributions, we propose to develop a kernel density deep neural network. Specifically, for face alignment, we adapt state-of-the-art hourglass neural network into a probabilistic neural network framework with landmark probability map as its output. The model is trained by maximizing the conditional log likelihood. To exploit the output probability map, we extend the model to multi-stage so that the logits map from the previous stage can feed into the next stage to progressively improve the landmark detection accuracy. Extensive experiments on benchmark datasets against state-of-the-art unconstrained deep learning method demonstrate that the proposed kernel density network achieves comparable or superior performance in terms of prediction accuracy. It further provides aleatoric uncertainty estimation in predictions. Lisha Chen, Hui Su |
ICCV | 2 |
| 2019 | Bayesian Graph Convolution LSTM for Skeleton Based Action RecognitionabstractWe propose a framework for recognizing human actions from skeleton data by modeling the underlying dynamic process that generates the motion pattern. We capture three major factors that contribute to the complexity of the motion pattern including spatial dependencies among body joints, temporal dependencies of body poses, and variation among subjects in action execution. We utilize graph convolution to extract structure-aware feature representation from pose data by exploiting the skeleton anatomy. Long short-term memory (LSTM) network is then used to capture the temporal dynamics of the data. Finally, the whole model is extended under the Bayesian framework to a probabilistic model in order to better capture the stochasticity and variation in the data. An adversarial prior is developed to regularize the model parameters to improve the generalization of the model. A Bayesian inference problem is formulated to solve the classification task. We demonstrate the benefit of this framework in several benchmark datasets with recognition under various generalization conditions. Rui Zhao 0015, Kang Wang 0002, Hui Su |
ICCV | 3 |
| 2019 | Machine Learning Accelerated Partition Search for Video EncodingabstractWith more complex partitioning structures in recent generations of video coding standards, the computation complexity of video encoder for partition block size search has been increasing drastically. To expedite the overall encoding process, it is desired to make faster partitioning decisions without much compression performance degradation. In this paper, we propose a multi-scale multi-stage machine learning(ML) based framework to accelerate partition block size search. The framework includes a collection of ML models, each dedicated to make a simple decision for a particular block size at a particular stage during the partitioning rate-distortion optimization(RDO) process. The ML models can predict whether the RD evaluation of certain partition block sizes can be skipped, saving unnecessary computation in the encoder. The proposed approach is implemented and tested on VP9 with the open source library libvpx. Significant encoding speed improvement has been observed with neglectable compression performance regression. The framework and methodology can be easily applied to other video codecs and implementations as well. Hui Su, Chi-Yo Tsai, Yunqing Wang, Yaowu Xu |
ICIP | 1 |
| 2019 | An Eye on the Storm: Uncovering Multi-Variate Relationships with a Science-Driven System For Interactive Analysis and Visualization; Motivating Machine-Learning Discoveries for Hurricane Rapid Intensity ChangesabstractThe paper discusses the hurricane intensity changes using machine learning technologies and data visualization. Svetla M. Hristova-Veleva, Bjorn Lambrigtsen, Hui Su, Jeffrey S. Reid, Saiprasanth Bhalachandran, Hua Leighton, Sundararaman Gopalakrishnan, Francisco J. Tapiador, P. Peggy Li, Brian W. Knosp, F. Joseph Turk, William Lee Poulsen, Quoc Vu, Ziad S. Haddad, Tsae-Pyng Shen, Bryan W. Stiles |
IGARSS | 3 |
| 2019 | Embodied Conversational AI Agents in a Multi-modal Multi-agent Competitive DialogueabstractIn a setting where two AI agents embodied as animated humanoid avatars are engaged in a conversation with one human and each other, we see two challenges. One, determination by the AI agents about which one of them is being addressed. Two, determination by the AI agents if they may/could/should speak at the end of a turn. In this work we bring these two challenges together and explore the participation of AI agents in multi-party conversations. Particularly, we show two embodied AI shopkeeper agents who sell similar items aiming to get the business of a user by competing with each other on the price. In this scenario, we solve the first challenge by using headpose (estimated by deep learning techniques) to determine who the user is talking to. For the second challenge we use deontic logic to model rules of a negotiation conversation. Rahul R. Divekar, Xiangyang Mou, Lisha Chen, Maíra Gatti de Bayser, Melina Alberio Guerra, Hui Su |
IJCAI | 6 |
| 2019 | Reagent: Converting Ordinary Webpages into Interactive Software AgentsabstractWe introduce Reagent, a technology that can be used in conjunction with automated speech recognition to allow users to query and manipulate ordinary webpages via speech and pointing. Reagent can be used out-of-the-box with third-party websites, as it requires neither special instrumentation from website developers nor special domain knowledge to capture semantically-meaningful mouse interactions with structured elements such as tables and plots. When it is unable to infer mappings between domain vocabulary and visible webpage content on its own, Reagent proactively seeks help by engaging in a voice-based interaction with the user. Matthew Peveler, Jeffrey O. Kephart, Hui Su |
IJCAI | 3 |
| 2019 | You Talkin' to Me? A Practical Attention-Aware Embodied Agent
Rahul R. Divekar, Jeffrey O. Kephart, Xiangyang Mou, Lisha Chen, Hui Su |
INTERACT (3) | 5 |
| 2019 | Variety matters: a new model for the wireless data market under sponsored data plansabstractIn this paper, we develop a new model to study the competition among Content Providers (CPs) under Sponsored Data Plans (SDPs). SDP is an emerging pricing model for the wireless data market where Internet Service Providers (ISPs) allow a CP to compensate the traffic volume of users when users access the contents of this CP. Studies have shown that SDPs create a triple-win situation, where users consume more contents and the revenue of both CPs and ISPs increases. Currently, a main concern of SDPs is on whether SDPs may bring about unfair competition among CPs. Studies have shown that big CPs have an advantage over small CPs. We observe that such conclusions are derived because in all previous models, traffic price is the only factor that affects user decisions. We argue that it is not precise. Nowadays, people conduct a large variety of activities online, and users have an intrinsic demand for a variety of contents. To reflect this, we for the first time characterize the variety demand as an intrinsic parameter of users, and integrate such variety into a new model to help us drive some novel insights into SDPs, especially the competition among CPs. Our model shows that variety matters for understanding SDPs more thoroughly and comprehensively. For example, under SDPs, the advantage of CPs with higher revenue will be significantly reduced if users have a greater love for variety. Overall, our new model leads to a set of completely new results and rectifies some past conclusions. Yi Zhao 0011, Hui Su, Liang Zhang 0042, Dan Wang 0002, Ke Xu 0002 |
IWQoS | 2 |
| 2019 | Deep Structured Prediction for Facial Landmark DetectionabstractExisting deep learning based facial landmark detection methods have achieved excellent performance. These methods, however, do not explicitly embed the structural dependencies among landmark points. They hence cannot preserve the geometric relationships between landmark points or generalize well to challenging conditions or unseen data. This paper proposes a method for deep structured facial landmark detection based on combining a deep Convolutional Network with a Conditional Random Field. We demonstrate its superior performance to existing state-of-the-art techniques in facial landmark detection, especially a better generalization ability on challenging datasets that include large pose and occlusion. Lisha Chen, Hui Su |
NeurIPS | 2 |
| 2019 | Machine Learning Accelerated Transform Search For AV1abstractAV1 is the state-of-the-art open and royalty-free video compression format that achieves significant bitrate savings over previous generation of video codecs. One of AV1's major improvement over its predecessor VP9 is the support of more diverse and flexible transform size and kernel selection. However, it also drastically increases the search space for transform unit rate-distortion optimization in AV1 encoders. Unlike conventional encoder speed features that are based on heuristics, we propose a machine learning (ML) based approach to accelerate the transform size and kernel search for AV1. The ML models use input features extracted from the prediction residue block such as standard deviation, correlation and energy distribution. The output of the models indicates the estimated likelihood of which transform size and kernel would be selected as the optimal choice. Based on the ML models, the encoder can prune out the transform size and kernel candidates that are unlikely to be selected and save unnecessary computation to compute their rate-distortion cost. The proposed approach is implemented and tested on the AV1 reference library libaom. The experimental results show that satisfactory encoding speed improvement can be achieved with extremely low compression performance loss. The framework and methodology can also be easily migrated to other video codecs and implementations. Hui Su, Alexander Bokov, Debargha Mukherjee, Yunqing Wang, Yue Chen 0040 |
PCS | 1 |
| 2018 | Interaction Challenges in AI Equipped Environments Built to Teach Foreign Languages Through Dialogue and Task-CompletionabstractAs cities around the world become more diverse in culture and language, there is a growing need for learning foreign languages. To further this excitement, we have built a human-scale, immersive room with a virtual AI agent that aids foreign language learning. Our system aids the language learning process through task-completion exercises using multi-modal dialogue. The Cognitive and Immersive Room (CIR) is developed as an immersive Chinese restaurant to teach Mandarin, but the interaction challenges and solutions can be reasonably generalized to other languages taught using similar techniques. As users interact with the immersive environment and the virtual AI agent, they face several user interaction challenges. These challenges arise from new learners' lack of proficiency in the foreign language. By studying user interactions in the CIR, we were able to articulate some of the interaction challenges. We have enhanced the AI agent, virtual environment, and the on-boarding process for new users to mitigate these challenges. The enhancements and the results which show that they were effective are discussed here. Rahul R. Divekar, Jaimie Drozdal, Yalun Zhou, Ziyi Song, Robert Rouhani, Rui Zhao 0015, Shuyue Zheng, Lilit Balagyozyan, Hui Su |
Conference on Designing Interactive Systems | 10 |
| 2018 | Towards Better Variational Encoder-Decoders in Seq2Seq TasksabstractVariational encoder-decoders have shown promising results in seq2seq tasks. However, the training process is known difficult to be controlled because latent variables tend to be ignored while decoding. In this paper, we thoroughly analyze the reason behind this training difficulty, compare different ways of alleviating it and propose a new framework that helps significantly improve the overall performance. Xiaoyu Shen 0001, Hui Su |
AAAI | 2 |
| 2018 | Improving Variational Encoder-Decoders in Dialogue GenerationabstractVariational encoder-decoders (VEDs) have shown promising results in dialogue generation. However, the latent variable distributions are usually approximated by a much simpler model than the powerful RNN structure used for encoding and decoding, yielding the KL-vanishing problem and inconsistent training objective. In this paper, we separate the training step into two phases: The first phase learns to autoencode discrete texts into continuous embeddings, from which the second phase learns to generalize latent representations by reconstructing the encoded embedding. In this case, latent variables are sampled by transforming Gaussian noise through multi-layer perceptrons and are trained with a separate VED model, which has the potential of realizing a much more flexible distribution. We compare our model with current popular models and the experiment demonstrates substantial improvement in both metric-based and human evaluations. Xiaoyu Shen 0001, Hui Su, Shuzi Niu, Vera Demberg |
AAAI | 2 |
| 2018 | Dialogue Generation With GANabstractThis paper presents a Generative Adversarial Network (GAN) to model multiturn dialogue generation, which trains a latent hierarchical recurrent encoder-decoder simultaneously with a discriminative classifier that make the prior approximate to the posterior. Experiments show that our model achieves better results. Hui Su, Xiaoyu Shen 0001, Pengwei Hu 0001, Wenjie Li 0002 |
AAAI | 1 |
| 2018 | Intra Block Copy for Screen Content in the Emerging AV1 Video CodecabstractScreen content coding plays an important role in many applications. To meet the growing demands of screen content coding, the emerging AV1 video codec incorporates several coding tools, which are specially designed for screen content utilizing its distinctive characteristics. Among these tools, the intra block copy utilizes the characteristic that repeating patterns frequently occur in screen content. This paper presents the technology of intra block copy in AV1. In particular, to efficiently search the predictor in the reconstructed regions of the current picture, AV1 uses the hash matching method at the encoder side. For the generation of hash table, a bottom-to-up manner is adopted to reduce the redundant computation and then decrease the encoding time. In addition, several constraints are involved to facilitate hardware design. Experimental results demonstrate that the intra block copy in AV1 can bring 27.1% bitrate saving for screen content. When compared with the non hash-based intra block copy, the hash-based method achieves 12.2% bitrate saving. Jiahao Li 0001, Hui Su, Alex Converse, Bin Li 0012, Roger Zhou, Bruce Lin, Jizheng Xu, Yan Lu 0001, Ruiqin Xiong |
DCC | 2 |
| 2018 | Nexus Network: Connecting the Preceding and the Following in Dialogue GenerationabstractSequence-to-Sequence (seq2seq) models have become overwhelmingly popular in building end-to-end trainable dialogue systems.Though highly efficient in learning the backbone of human-computer communications, they suffer from the problem of strongly favoring short generic responses.In this paper, we argue that a good response should smoothly connect both the preceding dialogue history and the following conversations.We strengthen this connection through mutual information maximization.To sidestep the nondifferentiability of discrete natural language tokens, we introduce an auxiliary continuous code space and map such code space to a learnable prior distribution for generation purpose.Experiments on two dialogue datasets validate the effectiveness of our model, where the generated responses are closely related to the dialogue context and lead to more interactive conversations.* Indicates equal contribution.X. Shen focuses on algorithm and H. Su is responsible for experiments. Xiaoyu Shen 0001, Hui Su, Wenjie Li 0002, Dietrich Klakow |
EMNLP | 2 |
| 2018 | An Immersive System with Multi-Modal Human-Computer InteractionabstractWe introduce an immersive system prototype that integrates face, gesture and speech recognition techniques to support multi-modal human-computer interaction capability. Embedded in an indoor room setting, a multi-camera system is developed to monitor the user facial behavior, body gesture and spatial location in the room. A server that fuses different sensor inputs in a time-sensitive manner so that our system knows who is doing what at where in real-time. When correlating with speech input, the system can better understand the user intention for interaction purpose. We evaluate the performance of core recognition techniques on both benchmark and self-collected datasets and demonstrate the benefit of the system in various use cases. Rui Zhao 0015, Kang Wang 0002, Rahul R. Divekar, Robert Rouhani, Hui Su |
FG | 5 |
| 2018 | An Overview of Core Coding Tools in the AV1 Video CodecabstractAV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz |
PCS | 9 |
| 2018 | A Cost-Effective Framework for Preference Elicitation and Aggregation
Zhibing Zhao, Haoming Li 0002, Jeffrey O. Kephart, Nicholas Mattei, Hui Su, Lirong Xia |
UAI | 6 |
| 2017 | Dow Jones Index is Driven Periodically by the Unemployment Rate During Economic Crisis and Non-economic Crisis Periods
Tong Cao, Sanqing Hu, Yuying Zhu 0005, Hui Su |
ICONIP (5) | 5 |
| 2017 | Wake-Sleep Variational Autoencoders for Language Modeling
Xiaoyu Shen 0001, Hui Su, Shuzi Niu, Dietrich Klakow |
ICONIP (1) | 2 |
| 2017 | Causality Analysis Between Soil of Different Depth Moisture and Precipitation in the United States
Hui Su, Sanqing Hu, Tong Cao, Yuying Zhu 0005 |
ICONIP (5) | 1 |
| 2017 | Identify Non-fatigue State to Fatigue State Using Causality Measure During Game Play
Yuying Zhu 0005, Yi-Ning Wu, Hui Su, Sanqing Hu, Tong Cao, Yu Cao 0002 |
ICONIP (4) | 3 |
| 2017 | All-Weather tropospheric 3D wind from microwave soundersabstractIn its 2007 “Decadal Survey” (DS) of earth science missions for NASA [1] the U.S. National Research Council (NRC) recommended that a Doppler wind lidar be developed for a three-dimensional tropospheric winds mission (“3D-Winds”). It is expected that the next DS, currently under way, will put additional emphasis on the still pressing need for wind measurements from space. The first DS also called for a geostationary microwave sounder (GMS) on a Precipitation and All-weather Temperature and Humidity (PATH) mission. Such a sounder, the Geostationary Synthetic Thinned Aperture Radiometer (GeoSTAR), has been developed at the Jet Propulsion Laboratory (JPL). The PATH mission has not yet been funded by NASA, but a low-cost subset of PATH, GeoStorm was proposed as a hosted payload on a commercial communications satellite. Both PATH and GeoStorm will obtain frequent (every 15 minutes of better) measurements of tropospheric water vapor profiles, and they can be used to derive atmospheric motion vector (AMV) wind profiles even in the presence of clouds. We report on simulation studies of such wind vectors derived from a GMS or possibly from a cluster of low-earth-orbiting (LEO) small satellites (e.g., CubeSats). Bjorn Lambrigtsen, F. Joseph Turk, Hui Su |
IGARSS | 3 |
| 2017 | DailyDialog: A Manually Labelled Multi-turn Dialogue DatasetabstractWe develop a high-quality multi-turn dialog dataset, DailyDialog, which is intriguing in several aspects. The language is human-written and less noisy. The dialogues in the dataset reflect our daily communication way and cover various topics about our daily life. We also manually label the developed dataset with communication intention and emotion information. Then, we evaluate existing approaches on DailyDialog dataset and hope it benefit the research field of dialog systems. The dataset is available on http://yanran.li/dailydialog Yanran Li, Hui Su, Xiaoyu Shen 0001, Wenjie Li 0002, Ziqiang Cao, Shuzi Niu |
IJCNLP(1) | 2 |
| 2015 | ESTRA: Incentivizing Storage Trading for Edge Caching in Mobile Content DeliveryabstractThe explosion of mobile content and usage imposes enormous pressures on mobile communication networks. To reduce content delivery latency and to ease the burden on network bottlenecks (e.g., backhaul networks), besides upgrading the infrastructures, it is promising to cache popular contents at the edge-storage on BSs (Base Stations), APs (Access Points) or other third-party devices associated with BSs and APs, which have been widely deployed in mobile networks. Then it is a challenge here to effectively match such demands of CPs (Content Providers) and the supplies of edge-storage owners because of the complexity caused by the two-fold matching requirements on both coverage and quantity combined with the multi-buyer multi-seller scenario and the divisibility of heterogeneous edge-storages. In this paper, we propose ESTRA (Edge Storage TRading Auction) mechanism to tackle such a challenge. By proposing a region-based model for edge-storage trading, we design a demand cover mechanism to transfer subscriber coverage demands into edge-storage bundle demands and design a truthful, weakly budget balanced and individually rational auction mechanism under the constraints of enabling multi-unit asks and bundle bids. Our theoretical analysis proves the economic robustness, and the simulation result shows that ESTRA achieves 74%-91% of the maximum social welfare and maintains the sustainability of the trading platform through a proper distribution of social welfare. Yifeng Zhong, Ke Xu 0002, Xiang-Yang Li 0001, Hui Su, Qingyang Xiao |
GLOBECOM | 4 |
| 2015 | TSP: A traffic sharing platform for mobile networksabstractIn mobile Internet era, wireless traffic has become a rare resource and there is no effective ways for users to share their unused traffic with each other. This paper introduces a system solution requiring no sophisticated hardware. An incentive mechanism is designed and implemented in a novel system named Traffic Sharing Platform (TSP) for mobile users, which can optimize network resource configuration and achieve Pareto optimality of the society. Simulation results show the TSP is available and the incentive mechanism is effective. Hui Su, Tong Li 0014, Ke Xu 0002, Shenglin Zhang, Xiaoliang Wang 0004 |
IWQoS | 1 |
| 2014 | Exploring the use of ENF for multimedia synchronizationabstractThe electric network frequency (ENF) signal can be captured in multimedia recordings due to electromagnetic influences from the power grid at the time of recording. Recent work has exploited the ENF signals for forensic applications, such as authenticating and detecting forgery of ENF-containing multimedia signals, and inferring their time and location of creation. In this paper, we explore a new potential of ENF signals for automatic synchronization of audio and video. The ENF signal as a time-varying random process can be used as a timing fingerprint of multimedia signals. Synchronization of audio and video recordings can be achieved by aligning their embedded ENF signals. We demonstrate the proposed scheme with two applications: multi-view video synchronization and synchronization of historical audio recordings. The experimental results show the ENF based synchronization approach is effective, and has the potential to solve problems that are intractable by other existing methods. Hui Su, Adi Hajj-Ahmad, Min Wu 0001, Douglas W. Oard |
ICASSP | 1 |
| 2014 | Exploiting rolling shutter for ENF signal extraction from videoabstractThe electric network frequency (ENF) signal can be embedded in multimedia recordings created in areas of electrical activities. Recent work has used the ENF signal for such applications as time stamp authentication and forgery detection. It is more challenging to extract ENF signals from video recordings than from audio recordings because of the low temporal sampling rate or frame rate of video cameras. The rolling shutter of CMOS image sensor can be exploited as it exposes a frame line by line, and the effective ENF sampling rate by treating each line as a signal sample can be increased. This scheme was shown to work well with static videos. This paper conducts a further study on the exploitation of the rolling shutter for extracting ENF traces from videos. The rolling shutter mechanism is modeled and analyzed using multirate signal processing theory. Challenging cases of videos with motions are examined, and solutions to extracting ENF from them are explored. Hui Su, Adi Hajj-Ahmad, Ravi Garg, Min Wu 0001 |
ICIP | 1 |
| 2013 | ENF analysis on recaptured audio recordingsabstractElectric Network Frequency (ENF) based forensic analysis is a promising tool for timestamp authentication and forgery detection in such multimedia recordings as audios and videos. ENF signal is embedded in an audio recording due to electromagnetic interference from the power lines. The time of creation of a multimedia recording can be determined by comparing the ENF signal embedded in the recording with a reference ENF database collected from the power grid. In this paper, we conduct a study of the effect of recapturing of audio recordings on the ENF embedding. We demonstrate that recaptured audio recordings pick up two ENF signals: the content ENF signal which is inherited from the original audio recording; and the recapturing ENF signal which is embedded from the recapturing process. Conventional ENF signal extraction techniques on such recordings may fail when the two ENF signals are at the same nominal value. A decorrelation algorithm is proposed to extract the content ENF signal and the recapturing ENF signal. The experimental results show the effectiveness of the proposed method in the estimation of both the ENF signals. Hui Su, Ravi Garg, Adi Hajj-Ahmad, Min Wu 0001 |
ICASSP | 1 |
| 2012 | Evaluating the quality of individual SIFT featuresabstractScale-Invariant Feature Transform (SIFT) is one of the most popular local image features that are widely used in computer vision, image processing and image retrieval. In this paper we study the relation between the SIFT descriptor and its matching accuracy. We propose a method to quantitatively assess the quality of a SIFT feature descriptor in terms of robustness and discriminability. This would enable us to gain a better understanding of the strength and limitations of SIFT in emerging applications of SIFT-based image hash, and also to improve matching accuracy and efficiency in applications such as object search. The experimental results demonstrate the effectiveness of the proposed method. Hui Su, Wei-Hong Chuang, Wenjun Lu, Min Wu 0001 |
ICIP | 1 |
| 2011 | Exploring compression effects for improved source camera identification using strongly compressed videoabstractThis paper presents a study of the video compression effect on source camera identification based on the Photo-Response Non-Uniformity (PRNU). Specifically, the reliability of different types of frames in a compressed video is first investigated, which shows quantitatively that I-frames are more reliable than P-frames for PRNU estimation. Motivated by this observation, a new mechanism for estimating the reference PRNU and two mechanisms for estimating the test-video PRNU are proposed to achieve higher accuracy with fewer frames used. Experiments are performed to validate the effectiveness of the proposed mechanisms. Wei-Hong Chuang, Hui Su, Min Wu 0001 |
ICIP | 2 |
| 2010 | Effects of communication style and time orientation on notification systems and anti-virus softwareabstractThe objectives of this study were (1) to investigate the effects of communication style (CS) and time orientation (TO) on people's perception of and proficiency in responding to notification systems and (2) to study the applications of these effects in the design of user notification in anti-virus software. Significant effects were found in the experiment; the results showed that users with a low-context CS can remember and make sense of the information provided by the notification system better than users with a high-context CS. Polychronic users perceive a lower level of interruption of the notification messages than monochronic users; polychronic users prefer rapid and accurate responses to the stimuli provided by the notification system, whereas monochronic users tend to avoid responding to the stimuli. Four sessions of focus group discussions were then carried out with users of different CS and TO, which focused on the virus attack warning provided by anti-virus software. The results should prove useful to aid designers in creating more effective and appealing ways to give notification or warnings to users. Ding-Long Huang, Pei-Luen Patrick Rau, Hui Su, Nan Tu |
Behav. Inf. Technol. | 3 |
| 2008 | Relooking at services science and services innovation
Jen-Yao Chung, Hui Su |
Serv. Oriented Comput. Appl. | 3 |
| 2007 | Test Case Generation for Collaborative Real-time Editing ToolsabstractEfficient testing of the collaborative and real-time aspects of collaborative real-time editing tools (CRETs) is extremely challenging. Changes in testing parameters and conflict resolution policies result in significant time and effort for modifying the corresponding test cases. In this paper, we propose a time-line diagram approach to visually model aspects of timing and collaborative conflicts useful for efficient generation of test cases. A specification language, ACDATE, is introduced to formally specify the corresponding test scenarios. We develop an algorithm that allows configuring test parameters and collaboration policies on the fly, automatically generates textual test cases corresponding to the timeline diagram, and creates corresponding test scripts in the ACDATE language. A prototype implementation shows good results, with automatically generated test cases consistent across from a visual and textual perspective. Wenping Xiao, Chang Yan Chi, Hui Su |
COMPSAC (1) | 5 |
| 2006 | The visual funding navigator: analysis of the NSF funding informationabstractThis paper presents an interactive visualization toolkit for navigating and analyzing the National Science Foundation (NSF) funding information. Our design builds upon an improved 2.5D treemap layout and the stacked graph to contribute customized techniques for visually navigating and interacting with the hierarchical data of NSF programs and proposals. Furthermore, an incremental layout method is adopted to handle information on a large scale. The improved treemap visualization will help to visually analyze the static funding related data and the stacked graph is utilized to analyze the time-series data. Through these visual analysis techniques, research trends of NSF, popular NSF programs are quickly identified. Shixia Liu, Nan Cao 0001, Hui Su |
CIKM | 4 |
| 2006 | Towards Facilitating Development of SOA Application with Design Metrics
Wei Zhao 0003, Ying Liu 0045, Hui Su |
ICSOC | 4 |
| 2005 | Adaptive predict based on fading compensation for lifting-based motion compensated temporal filteringabstractA lifting implementation of the discrete wavelet transform applied along motion trajectories has recently gained a lot of attention in the video community as strong candidates in incoming scalable video coders. We generalize the coding scheme for classical lifting-based motion compensation temporal filtering and permit the codec to choose adaptively between the original reference frames and new fading-compensated reference frames to predict residuals while maintaining the invertibility of the inter-frame transform. Experimental results show that the proposed algorithm not only significantly improves subjective visual quality of the temporal low-pass frames, but also has 0.15-0.3 dB gain in PSNR performance compared with the normal (5, 3) lifting schemes. Li Song 0001, Hongkai Xiong, Jizheng Xu, Feng Wu 0001, Hui Su |
ICASSP (2) | 5 |
| 2003 | Multimodal Menu Interface for Mobile Web Browsing
Paul P. Maglio, Hui Su |
INTERACT | 3 |
| 2001 | Chinese input with keyboard and eye-tracking: an anatomical studyabstractChinese input presents unique challenges to the field of human computer interaction. This study provides an anatomical analysis of today's standard Chinese input process, which is based on pinyin, a phonetic spelling system in Roman characters. Through a combination of human performance modeling and experimentation, our study decomposed the Chinese input process into sub-tasks and found that choice reaction time and numeric keying, two component resulted from the large number of homophones in Chinese, were the major usability bottlenecks. Choice reaction alone took 36% of the total input time in our experiment. Numeric keying for multiple candidates selection tends to take the user's attention away from the computer visual screen. We designed and implemented the EASE (Eye Assisted Selection and Entry) system to help maintaining complete touch-typing experience without diverting visual (spacebar) and implicit eye-tracking to replace the numeric keystrokes. Our experiment showed that such a system could indeed work, even with today's imperfecteye-tracking technology. Shumin Zhai, Hui Su |
CHI | 3 |
| 2001 | Detectability and Comprehensibility Study on Audio Hyperlinking Methods
Qian Ying Wang, MoWei Shen, RenDe Shui, Hui Su |
INTERACT | 4 |
| 2000 | Automatic text extraction from color image
Hui Su, Chang Yan Chi |
VCIP | 2 |
| 1997 | A New Method for Segmenting Unconstrained Handwritten Numeral StringabstractA new segmentation method for segmenting an unconstrained handwritten numeral string with an unknown number of digits is proposed, which is mainly based on recognition-based segmentation, combined with the dissection and holistic methods. The authors proposed an approach that gives a simple, effective way of selecting candidate segmentation positions by finding the start points and end points of ligatures. The approach is mainly based on the variety of upper and lower contours, combined with the vertical projection value. They also provide a method of confirming the correct segmenting points and deciding which digit the connecting stroke belongs to by using dynamic programming. Aiming at decreasing the work time and complexity of segmentation, they provide a dissection method as well as the holistic method for segmenting some special sub-images. The experimental results are given. Hui Su, Shaowei Xia |
ICDAR | 2 |
| 1997 | A Fault-Tolerant Chinese Check Recognition SystemabstractIn this paper, a complete fault-tolerant check recognition system is proposed which has no check substitution error under the secret code verification. The fault-tolerant recognition method proposed in this paper creates all possible candidates for verification under the limited fault-tolerant rate, and with three classifiers of high isolated digit recognition rate, the system can always find out the correct recognition results of checks if there exist the correct labels of all the digits. Since the three classifiers are designed independently by different methods and they extract different features of handwritten digits, they can compensate each other when confusing digits are met. The segmentation stage combines the three most popular strategies, and gives out a way for segmenting unconstrained handwritten numeral strings on Chinese checks. Hui Su, Shaowei Xia |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1995 | Hierarchical neural network for recognizing hand-written characters in engineering drawingsabstractThe similarity degree of hand-written characters in engineering drawings is analyzed. On the basis of the results, a hierarchical neural network is proposed for recognizing hand-written characters in engineering drawings. In a hierarchical neural network the memorizing and recognizing procedure is distributed into several subnetworks. Experimental results are also given. Hui Su, Xinyou Li, Shaowei Xia |
ICDAR | 1 |