EDBT 2026 Demo / reviewers in the wild / expert
Rize Jin
dblp:41/8607
· DBLP profile ↗
23ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0001-5537-0744ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information ExtractionabstractMultimodal information extraction (MIE) constitutes a set of essential tasks aimed at extracting structural information from Web texts with integrating images, to facilitate the structural construction of Web-based semantic knowledge. To address the expanding category set including newly emerging entity types or relations on websites, prior research proposed the zero-shot MIE (ZS-MIE) task which aims to extract unseen structural knowledge with textual and visual modalities. However, the ZS-MIE models are limited to recognizing the samples that fall within the unseen category set, and they struggle to deal with real-world scenarios that encompass both seen and unseen categories. The shortcomings of existing methods can be ascribed to two main aspects. On one hand, these methods construct representations of samples and categories within Euclidean space, failing to capture the hierarchical semantic relationships between the two modalities within a sample and their corresponding category prototypes. On the other hand, there is a notable gap in the distribution of semantic similarity between seen and unseen category sets, which impacts the generative capability of the ZS-MIE models. To overcome the above disadvantages, we delve into the generalized zero-shot MIE (GZS-MIE) task and propose the hyperbolic multimodal generative representation learning framework (HMGRL). The variational information bottleneck and autoencoder networks are reconstructed with hyperbolic space for modeling the multi-level hierarchical semantic correlations among samples and prototypes. Furthermore, the proposed model is trained with the unseen samples generated by the decoder, and we introduce the semantic similarity distribution alignment loss to enhance the model's generalization performance. Experimental evaluations on two benchmark datasets underscore the superiority of HMGRL compared to existing baseline methods. Baohang Zhou, Kehui Song, Rize Jin, Yu Zhao 0043, Xuhui Sui, Xinying Qian, Xingyue Guo, Ying Zhang 0015 |
WWW | 3 |
| 2026 | A duet of perception and reasoning: CLIP and LLM brainstorming for scene text recognition
Zeguang Jia, Kehui Song, Zhilan Wang, Rize Jin |
Neurocomputing | 6 |
| 2025 | Mind Map-guided Meta-prompting for ADHD Intervention with Large Language Models
Xuguang Qiu, Kehui Song, Rize Jin |
ICONIP (2) | 4 |
| 2025 | Rosetta: Enhancing Neural Machine Translation through Multilingual-PLMs and Semantic ConstraintsabstractPreserving semantic equivalence between source text and its translation remains a formidable challenge in machine translation. Despite recent advancements, neural machine translation (NMT) models often fall short due to their over-reliance on word-level alignment, a limitation exacerbated by the use of cross-entropy loss. This paper introduces Rosetta, a novel framework that revisits the semantic awareness of NMT Systems. Our approach builds upon the Transformer architecture, incorporating stochastic layer selection, multi-branch attention, and group fusion mechanisms to facilitate a more nuanced understanding of context and meaning. Central to our framework is the utilization of multilingual pre-trained language models (PLMs) to reinforce semantic consistency between the source text and its translation in the semantic space. We evaluate our approach on the IWSLT’14 translation dataset. Our model achieves state-of-the-art performance on the German-English translation task, surpassing existing models with a BLEU score of 39.12. To assess the semantic equivalence of translations, we employ GPT-4 as an independent evaluator. The results demonstrate that our approach surpasses literal translation, excelling in preserving semantic similarity. Pengfei Pi, Rize Jin, Tae-Sun Chung |
IJCNN | 2 |
| 2025 | MaCSE: Multi-Agent Ranking Distillation for Contrastive Learning of Sentence EmbeddingsabstractSentence embedding models are typically trained using the Contrastive Learning (CL) method, which works by pulling similar semantics closer and pushing dissimilar ones away. Recent studies have shown that utilizing a multi-teacher ranking distillation approach, which assigns fine-grained rankings to sentences, enables the generation of smoother sentence similarity representations and results in higher-quality sentence embeddings. However, the effectiveness of distillation may be limited by the capacity of the student model. A simple student model with fewer parameters may struggle to approximate a highly complex teacher model, potentially leading to overfitting on certain datasets or specific aspects of the task. To address this, we propose MaCSE, a multi-agent ranking distillation framework that dynamically selects and optimizes teacher model contributions across training stages. MaCSE employs a Centralized Training with Decentralized Execution (CTDE) paradigm, enabling collaborative agent interactions to adaptively adjust teacher fusion weights based on training dynamics. Experimental results on Semantic Textual Similarity and transfer tasks demonstrate that MaCSE outperforms most existing baselines and even rivals methods using large language models for sentence representation. Our implementation is available at GitHub1. Zekai Zhi, Zhilan Wang, Rize Jin, Kehui Song, Da-Jung Cho |
IJCNN | 3 |
| 2025 | From Chain to Loop: Improving Reasoning Capability in Small Language Models via Loop-of-Thought
Mingxin Ji, Kehui Song, Rize Jin, Xuguang Qiu |
NLPCC (1) | 3 |
| 2025 | FaceDisentGAN: Disentangled facial editing with targeted semantic alignment
Meng Xu 0024, Prince Hamandawana, Zekang Chen, Rize Jin, Tae-Sun Chung |
Neurocomputing | 5 |
| 2024 | Multi-Channel Spatio-Temporal Transformer for Sign Language ProductionabstractThe task of Sign Language Production (SLP) in machine learning involves converting text-based spoken language into corresponding sign language expressions. Sign language conveys meaning through the continuous movement of multiple articulators, including manual and non-manual channels. However, most current Transformer-based SLP models convert these multi-channel sign poses into a unified feature representation, ignoring the inherent structural correlations between channels. This paper introduces a novel approach called MCST-Transformer for skeletal sign language production. It employs multi-channel spatial attention to capture correlations across various channels within each frame, and temporal attention to learn sequential dependencies for each channel over time. Additionally, the paper explores and experiments with multiple fusion techniques to combine the spatial and temporal representations into naturalistic sign sequences. To validate the effectiveness of the proposed MCST-Transformer model and its constituent components, extensive experiments were conducted on two benchmark sign language datasets from diverse cultures. The results demonstrate that this new approach outperforms state-of-the-art models on both datasets. Rize Jin, Tae-Sun Chung |
LREC/COLING | 2 |
| 2024 | Reinforced Multi-teacher Knowledge Distillation for Unsupervised Sentence Representation
Rize Jin, Shibo Qi |
ICANN (7) | 2 |
| 2024 | GRNet: a graph reasoning network for enhanced multi-modal learning in scene text recognitionabstractAbstract Recent advancements in scene text recognition have predominantly focused on leveraging textual semantics. However, an over-reliance on linguistic priors can impede a model’s ability to handle irregular text scenes, including non-standard word usage, occlusions, severe distortions, or stretching. The key challenges lie in effectively localizing occlusions, perceiving multi-scale text, and inferring text based on scene context. To address these challenges and enhance visual capabilities, we introduce the Graph Reasoning Model (GRM). The GRM employs a novel feature fusion method to align spatial context information across different scales, beginning with a feature aggregation stage that extracts rich spatial contextual information from various feature maps. Visual reasoning representations are then obtained through graph convolution. We integrate the GRM module with a language model to form a two-stream architecture called GRNet. This architecture combines pure visual predictions with joint visual-linguistic predictions to produce the final recognition results. Additionally, we propose a dynamic iteration refinement for the language model to prevent over-correction of prediction results, ensuring a balanced contribution from both visual and linguistic cues. Extensive experiments demonstrate that GRNet achieves state-of-the-art average recognition accuracy across six mainstream benchmarks. These results highlight the efficacy of our multi-modal approach in scene text recognition, particularly in challenging scenarios where visual reasoning plays a crucial role. Zeguang Jia, Rize Jin |
Comput. J. | 3 |
| 2024 | Malware Family Prediction with an Awareness of Label UncertaintyabstractAbstract Malware family prediction has been mainly formulated as a multiclass classification to predict one malware family. This approach suffers from label uncertainty, which can mislead malware analysts. To render malware prediction less susceptible to uncertainty, malware family prediction, which entails predicting one or more families, is performed in this study. In this regard, an encoder–decoder malware family prediction model, EnDePMal, with label uncertainty awareness, is proposed. EnDePMal aims to predict all malware families related to samples and preserve their priorities. It comprises a residual neural network-based encoder and a long short-term memory-based decoder with an attention mechanism. The model uses a sequence of malware family names, but not a family name, as a label. Once a visualized malware image is input into EnDePMal, its encoder extracts the important features from the image. Subsequently, its decoder generates family names, where the attention mechanism allows it to focus on relevant features by attending to the encoder’s output. Experimental results show that EnDePMal can predict 77.64% of malware family sequences that preserve their priorities. Moreover, it achieves an accuracy of 93.49% and an F1-score of 0.9282 for malware families with the highest priority, rendering it comparable to the typical multiclass classification model. Joon-Young Paik, Rize Jin |
Comput. J. | 2 |
| 2024 | Attentional bias for hands: Cascade dual-decoder transformer for sign language productionabstractAbstract Sign Language Production (SLP) refers to the task of translating textural forms of spoken language into corresponding sign language expressions. Sign languages convey meaning by means of multiple asynchronous articulators, including manual and non‐manual information channels. Recent deep learning‐based SLP models directly generate the full‐articulatory sign sequence from the text input in an end‐to‐end manner. However, these models largely down weight the importance of subtle differences in the manual articulation due to the effect of regression to the mean. To explore these neglected aspects, an efficient cascade dual‐decoder Transformer (CasDual‐Transformer) for SLP is proposed to learn, successively, two mappings SLP hand : Text → Hand pose and SLP sign : Text → Sign pose , utilising an attention‐based alignment module that fuses the hand and sign features from previous time steps to predict more expressive sign pose at the current time step. In addition, to provide more efficacious guidance, a novel spatio‐temporal loss to penalise shape dissimilarity and temporal distortions of produced sequences is introduced. Experimental studies are performed on two benchmark sign language datasets from distinct cultures to verify the performance of the proposed model. Both quantitative and qualitative results show that the authors’ model demonstrates competitive performance compared to state‐of‐the‐art models, and in some cases, achieves considerable improvements over them. Rize Jin, Tae-Sun Chung |
IET Comput. Vis. | 2 |
| 2024 | Channel and Spatial Enhancement Network for human parsing
Kunliang Liu, Rize Jin, Wonjun Hwang |
Image Vis. Comput. | 2 |
| 2023 | Unsupervised Contrastive Learning of Sentence Embeddings Through Optimized Sample Construction and Knowledge Distillation
Rize Jin, Joon-Young Paik, Tae-Sun Chung |
PRICAI (2) | 2 |
| 2023 | Neural Machine Translation with an Awareness of Semantic Similarity
Rize Jin, Joon-Young Paik, Tae-Sun Chung |
PRICAI (2) | 2 |
| 2022 | Malware classification using a byte-granularity feature based on structural entropyabstractAbstract Rapidly evolving malware has become a major cybersecurity threat. Several feature‐engineering techniques have been proposed to defend against malware attacks. An entropy is a typical indicator used in identifying malware. Structural entropy is a sequence of entropy values where an entropy of a segment is calculated by the equation of the entropy itself. However, entropy‐based features are likely to be abstract and miss important information. This article proposes a feature engineering technique that involves the concept of structural entropy. This technique allows every segment to be represented as 256 entropy values for every byte value, but not as an entropy value. Our research, fine‐granularity structural entropy (FiG_SE), incorporates global patterns across all segments, local patterns across adjacent segments, and internal patterns within the segments. To extract higher‐level characteristics from our entropy feature, we use a convolutional neural network (CNN) architecture because it is effective for extracting local and global patterns, and especially for shift‐invariant patterns. Our malware classification based on CNN with the proposed feature outperforms the previous classification methods that use byte streams, entropy streams, and structural‐entropy‐based streams as inputs. Moreover, our research combined with CNN is highly resilient to obfuscation techniques and is also well suited to malware detection. Joon-Young Paik, Rize Jin, Eun-Sun Cho |
Comput. Intell. | 2 |
| 2020 | Managing Massive Amounts of Small Files in All-Flash StorageabstractAll-flash array is a popular memory device available for use in modern high-performance storage systems. Compared with other types of devices such as DRAM, NVRAM, and EEPROM, flash array combines the best features: shock resistance, low cost, low power consumption, and fast access. Moreover, the ever-increasing density of flash memory has led to a dramatic increase in the capacity, which allows the storage of large volume of data. However, flash memory is not optimal for managing a large number of small files because: 1. the small and random write operation is inefficient in flash memory; 2. massive metadata information occupies a significant portion of the namespace, which is relatively limited or scarce in big data storage systems. This paper introduces a novel approach, hash partitioning-based file compaction (HFC), to improve the efficiency of storing and accessing small files in all-flash storage systems. HFC consists of a file compaction tool and an access interface. The compaction tool merges a group (usually a directory) of small files into a set of "big files" to reduce the metadata required to be maintained in the on-chip memory. The data locality and tree structure of those small files are preserved. The access interface is designed to provide transparent access to the small files in the HFC big files. Experimental results confirm that the proposed method significantly enhances the efficiency of managing massive amounts of small files in flash memory in terms of namespace usage and access speed. Rize Jin, Joon-Young Paik, Yenewondim Biadgie, Yunbo Rao, Tae-Sun Chung |
COMPSAC | 1 |
| 2019 | Toward Machine Learning Based Analyses on Compressed FirmwareabstractAs Internet of Things (IoT) applications are getting attention these days, the importance of firmware security is also growing. However, it is not straightforward to analyze the bugs or vulnerabilities that reside in firmware. One of the major challenges is to detect information about hardware architectures of compressed firmware. Traditional analysis tools make use of static signatures embedded in the compressed binary code of firmware. However, signature extraction needs the careful elaboration of experts, and it is not always even possible. In this paper, we introduce our experience in analyzing the hardware information of compressed firmware. Since it is not possible to use the semantic information of compressed binary code, we adopt machine learning technologies for this purpose. Despite various difficulties, we have positive experimental results. Seoksu Lee, Joon-Young Paik, Rize Jin, Eun-Sun Cho |
COMPSAC (2) | 3 |
| 2018 | A Storage-level Detection Mechanism against Crypto-RansomwareabstractRansomware represents a significant threat to both individuals and organizations. Moreover, the emergence of ransomware that exploits kernel vulnerabilities poses a serious detection challenge. In this paper, we propose a novel ransomware detection mechanism at a storage device, especially a flash-based storage device. To this end, we design a new buffer management policy that allows our detector to identify ransomware behaviors. Our mechanism detects a realistic ransomware sample with little negative impacts on the hit ratios of the buffers internally located in a storage device. Joon-Young Paik, Joong-Hyun Choi, Rize Jin, Eun-Sun Cho |
CCS | 3 |
| 2018 | Pose Specification Based Online Person Identification
Rize Jin, Guanghao Jin |
PKAW | 3 |
| 2018 | Improving Generative Adversarial Networks with Adaptive Control LearningabstractGenerative adversarial networks (GANs) are well known both for being unstable to train and for the problem of mode collapse, particularly when trained on data collections containing a diverse set of visual objects. This study introduces an adaptive hyper-parameter learning procedure for GANs as an alternative to the existing static approach. The proposed procedure is designed to mitigate the impact of instability and saturation in the original by dynamically adjusting the ratio of the training steps of both the generator and discriminator. To accomplish this, we track and analyze stable training curves of relatively narrow datasets and use them as the target fitting lines when training more diverse data collections. Experimental results show that the proposed model improves the stability and generates more realistic images. Rize Jin, Kyung-Ah Sohn 0001, Joon-Young Paik, Tae-Sun Chung |
VCIP | 2 |
| 2015 | A privacy-aware monitoring algorithm for moving k-nearest neighbor queries in road networks
Hyung-Ju Cho, Se Jin Kwon, Rize Jin, Tae-Sun Chung |
Distributed Parallel Databases | 3 |
| 2015 | A collaborative approach to moving k-nearest neighbor queries in directed and dynamic road networks
Hyung-Ju Cho, Rize Jin, Tae-Sun Chung |
Pervasive Mob. Comput. | 2 |