Junjie Shan

dblp:167/9032 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning
abstract
Foundation models that bridge vision and language have made significant progress. While they have inspired many life-enriching applications, their potential for abuse in creating new threats remains largely unexplored. In this paper, we reveal that vision-language models (VLMs) can be weaponized to enhance gradient inversion attacks (GIAs) in federated learning (FL), where an FL server attempts to reconstruct private data samples from gradients shared by victim clients. Despite recent advances, existing GIAs struggle to reconstruct high-resolution images when the victim has a large local data batch. One promising direction is to focus reconstruction on valuable samples rather than the entire batch, but current methods lack the flexibility to target specific data of interest. To address this gap, we propose Geminio, the first approach to transform GIAs into semantically meaningful, targeted attacks. It enables a brand new privacy attack experience: attackers can describe, in natural language, the data they consider valuable, and Geminio will prioritize reconstruction to focus on those high-value samples. This is achieved by leveraging a pretrained VLM to guide the optimization of a malicious global model that, when shared with and optimized by a victim, retains only gradients of samples that match the attacker-specified query. Geminio can be launched at any FL round and has no impact on normal training (i.e., the FL server can steal clients' data while still producing a high-utility ML model as in benign scenarios). Extensive experiments demonstrate its effectiveness in pinpointing and reconstructing targeted samples, with high success rates across complex datasets and large batch sizes with resilience against defenses.
Junjie Shan, Jialin Lu, Siu-Ming Yiu, Ka-Ho Chow 0001
ICCV1
2024 StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
abstract
Shihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shihan Dou, Yan Liu 0002, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang 0001, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
ACL (1)6
2024 SaProt: Protein Language Modeling with Structure-aware Vocabulary
abstract
Large-scale protein language models (PLMs), such as the ESM family, have achieved remarkable performance in various downstream tasks related to protein structure and function by undergoing unsupervised training on residue sequences. They have become essential tools for researchers and practitioners in biology. However, a limitation of vanilla PLMs is their lack of explicit consideration for protein structure information, which suggests the potential for further improvement. Motivated by this, we introduce the concept of a ``structure-aware vocabulary" that integrates residue tokens with structure tokens. The structure tokens are derived by encoding the 3D structure of proteins using Foldseek. We then propose SaProt, a large-scale general-purpose PLM trained on an extensive dataset comprising approximately 40 million protein sequences and structures. Through extensive evaluation, our SaProt model surpasses well-established and renowned baselines across 10 significant downstream tasks, demonstrating its exceptional capacity and broad applicability. We have made the code, pre-trained model, and all relevant materials available at https://github.com/westlake-repl/SaProt.
Jin Su, Chenchen Han, Junjie Shan, Xibin Zhou, Fajie Yuan
ICLR4
2024 Matching of Japanese Listening Test Dialogues and Anime Scene Dialogues based on Zero-shot Attribute Classification
abstract
Zero-shot Classification methods do not require extra training processes, but their Classification effectiveness differs from the input label set expected to be classified. This paper proposes a method to find label sets with the best Classification effectiveness and classify and match dialogues with them by zero-shot Classification methods. We investigated two Classification methods in this paper: (1) text embedding-based cosine similarity and (2) end-to-end pre-trained zero-shot model. We collected 250 listening test dialogues from each level of the past Japanese Language Proficiency Test (JLPT) and manually classified them by three attributes: (1) dialogue location, (2) speaker’s relationship, and (3) dialogue style. We used these listening test dialogues to test the effectiveness of zero-shot Classification under different input label sets. After comparing 212 label sets by RMSE (Root Mean Square Error), we identified seven label sets with the best Classification effectiveness. In the evaluation experiment, 314,930 anime scenes were classified with the seven label sets. We matched anime dialogue scenes and past listening test dialogues with their label under different zero-shot Classification methods and different numbers of attributes. We calculated the word cover rate and the text similarity between matched anime dialogue scenes and listening dialogues. The result shows that, compared with the random-sampling baseline, the proposed method using text embedding-based cosine similarity can reduce the number of anime scene candidates to 18.7% and result in a 0.82% increase in word cover rate and a 0.0285 increase in text similarity. In contrast, the end-to-end zero-shot model could reduce anime scene candidates to 15.2% and increase the word cover rate and text similarity with 2.13% and 0.0054, respectively.
Yangdi Ni, Junjie Shan, Yoko Nishihara
KES2
2024 Estimating Imagined Colors from Different Music Genres with Eye-Tracking
abstract
Music brings different feelings with different genres and types, and people would imagine different colors from those feelings. This has brought about a long-standing interest in the correspondence between colors and music. In this study, we propose a method that uses the eye-tracking device to record and estimate the colors that users imagined from listening to music. By showing randomly generated color tiles to music-listening users and asking them to find the colors they imagined from the music, we record the colors that users have gazed at through an eye-tracking system and estimate the colors that users have imagined with a clustering analysis algorithm. With this method, we can intuitively compare and analyze the correlation between music and color that is perceived by people without involving complex audio signal processing or expert music theories. Through an experiment with 10 participants of Japanese university students using a total of 18 songs in six categories of male vocals and female vocals of Rock, Ballad, and Pop music, we found that people would imagine similar colors when listening to the same music genres. Among the users’ gazed color records from the same music genre, the genre of the male-vocal rock had the most concentrated imagine-color, while the imagined color of the female-vocal ballad was the most dispersed.
Junjie Shan, Taijiro Nishizawa, Yoko Nishihara
KES1
2024 Short-form Video Needs Long-term Interests: An Industrial Solution for Serving Large User Sequence Models
abstract
Sequential models are invaluable for powering personalized recommendation systems. In the context of short-form video (SFV) feeds, where user behavior history is typically longer, systems must be able to understand users’ long-term interests. However, deploying large sequence models to extensive web-scale applications faces challenges due to high serving cost. To address this, we propose an industrial framework designed for efficiently serving large user sequence models. Specifically, the proposed infrastructure decouples serving of the user sequence model and the main recommendation model, with the user sequence model being served offline (asynchronously) with periodical refresh. The proposed infrastructure is also model-agnostic; thus, it can be used to support any type of user sequence models (even LLMs) with controllable costs. Empirical results show that large user models deployed with our framework significantly and consistently enhance the quality of the main recommendation model with minimal serving costs increase.
Yuening Li, Diego Uribe, Jiaxi Tang, Qingyun Liu 0003, Junjie Shan, Ben Most, Kaushik Kalyan, Shuchao Bi, Xinyang Yi, Lichan Hong, Ed H. Chi, Liang Liu 0017
RecSys6
2023 Multitask Ranking System for Immersive Feed and No More Clicks: A Case Study of Short-Form Video Recommendation
abstract
In recent years, social media users spend significant amount of time on Short-Form Video (SFV) platforms. Its success in creating an immersive viewership experience is not only from the content, but also due to its unique UI innovation: instead of providing choices for users to click, SFV platforms actively recommend content to users to watch one at a time. In this paper, we highlight unique challenges rooted from such UI changes for SFV recommendation system design. Firstly, there is yet much unexplored for sources of system biases under the new UI, as there are no clicks nor the common click-based position biases. Additionally, when training multiple types of user activities, positive labels for activities like sharing and commenting can be much sparser and more skewed than traditional click-based recommendation systems, as the latter can filter non-click impressions when generating "post-click" activities.
Qingyun Liu 0003, Zhe Zhao 0001, Liang Liu 0017, Junjie Shan, Yuening Li, Shuchao Bi, Lichan Hong, Ed H. Chi
CIKM5
2023 Sentiment Analysis of User Reviews Transition in Multimedia Franchise
abstract
A Multimedia franchise refers to developing a single production across various media such as comics, anime, and film. The media types are diverse, including comics, anime, movies, games, novels, and stage. People who enjoy a single production in one medium may enjoy the same production in another. A single production can be enjoyed in multiple media. However, evaluations for production are different among all franchised media. A production franchised in a medium obtains high evaluations, while the same production franchised in another medium may obtain low evaluations. Production evaluations in a medium are often conducted by comparing to previous evaluations in another medium. Therefore, evaluating a production franchised across various media may be related to the order of multimedia franchises. By analyzing the order of multimedia franchises and the evaluation of production together, it may be possible to gain knowledge on obtaining high evaluations in the multimedia franchise. We assume a correlation exists between the order of multimedia franchises and the production evaluation. In this paper, we group productions according to the order of multimedia franchises and conduct sentiment analysis on reviews to find the tendency of evaluation in each media within the group. For the analysis, we create a database that shows the progression of multimedia franchises for each production. The database includes the presence/absence of 10 types of multimedia franchises for 101 productions and the number of years. We choose four orders from the database's most common multimedia franchise types. We then collect reviews of production for each medium and conduct sentiment analysis. We calculate the rate of positive sentences of review of the production for each medium. Then, we analyze the relationship between multimedia franchise order and the polarity of reviews. Analysis results showed that the rate of positive sentences was highest for “stage,” followed by “comic,” and the lowest for “anime.” Furthermore, the transition from “anime” to “comic” increased by about 9% in positive sentences of reviews, while the transition from “comic” to “stage” decreased by about 4% in positive sentences of reviews. In the future, we plan to analyze the content of reviews and further explore the relationship between the multimedia franchise order and the reviews’ polarity.
Yuna Fujii, Junjie Shan, Yihong Han, Yoko Nishihara
KES2
2023 Gitor: Scalable Code Clone Detection by Building Global Sample Graph
abstract
Code clone detection is about finding out similar code fragments, which has drawn much attention in software engineering since it is important for software maintenance and evolution. Researchers have proposed many techniques and tools for source code clone detection, but current detection methods concentrate on analyzing or processing code samples individually without exploring the underlying connections among code samples.
Junjie Shan, Shihan Dou, Yueming Wu 0001, Hairu Wu, Yang Liu 0003
ESEC/SIGSOFT FSE1
2022 Decorrelate Irrelevant, Purify Relevant: Overcome Textual Spurious Correlations from a Feature Perspective
abstract
Natural language understanding (NLU) models tend to rely on spurious correlations (i.e., dataset bias) to achieve high performance on in-distribution datasets but poor performance on out-of-distribution ones. Most of the existing debiasing methods often identify and weaken these samples with biased features (i.e., superficial surface features that cause such spurious correlations). However, down-weighting these samples obstructs the model in learning from the non-biased parts of these samples. To tackle this challenge, in this paper, we propose to eliminate spurious correlations in a fine-grained manner from a feature space perspective. Specifically, we introduce Random Fourier Features and weighted re-sampling to decorrelate the dependencies between features to mitigate spurious correlations. After obtaining decorrelated features, we further design a mutual-information-based method to purify them, which forces the model to learn features that are more relevant to tasks. Extensive experiments on two well-studied NLU tasks demonstrate that our method is superior to other comparative approaches.
Shihan Dou, Songyang Gao, Junjie Shan, Qi Zhang 0001, Yueming Wu 0001, Xuanjing Huang 0001
COLING5
2022 TreeCen: Building Tree Graph for Scalable Semantic Code Clone Detection
abstract
Code clone detection is an important research problem that has attracted wide attention in software engineering. Many methods have been proposed for detecting code clone, among which text-based and token-based approaches are scalable but lack consideration of code semantics, thus resulting in the inability to detect semantic code clones. Methods based on intermediate representations of codes can solve the problem of semantic code clone detection. However, graph-based methods are not practicable due to code compilation, and existing tree-based approaches are limited by the scale of trees for scalable code clone detection.
Deqing Zou, Junru Peng, Yueming Wu 0001, Junjie Shan, Hai Jin 0001
ASE5
2018 A System for Japanese Listening Training Support with Watching Japanese Anime Scenes
abstract
As the widespread of Japanese entertainments and pop-cultures around the world, more and more people choose to learn Japanese as a foreign language (JFL). With the increasing of JFL learners, according to the survey of Japan Foundation, traditional materials are becoming inadequate for learners’ demands. For foreign language learning, the communication skills in listening and speaking would be more important and difficult at learning process, so are the JFL learners [1]. This research proposed a method using Japanese animation, which is called "Anime" in the world, as a learning support material for training JFL learners’ Japanese listening skill. Anime has a huge amount of scenes in which the characters speak usually with the standard and clear pronunciations. These dialogue scenes could be useful examples of Japanese conversation for JFL learners to practice their speaking and listening. The proposed system classifies the dialogue scenes in Anime depending on the dialogues’ degrees of difficulty. We firstly subdivided the dialogue scenes from each Anime episode. The system analyzes Japanese words and expressions used in the Anime dialogues and compares those with the scripts of listening tests from the previous Japanese Language Proficiency Tests (JLPT), like TOEIC or TOEFL for English. The system uses the cosine similarity of Japanese words and expressions frequency to estimate the corresponding degree level for each Anime scene and provides the appropriate scenes to JFL learners for listening training. The experimental results showed that the proposed system worked successfully and improved the correct answer rate of JFL learners more than 10% in listening tests.
Junjie Shan, Yoko Nishihara, Ryosuke Yamanishi
KES1
2017 Analysis of Dialogues Difficulty in Anime Comparing with JLPT Listening Tests
abstract
More and more people choose to learn Japanese as a foreign language (JFL) in the world. The traditional teaching materials have focused on training for the skills of writing and reading in Japanese. However, it seems that the JFL learners would like to train their speaking and listening rather than writing and reading because the speaking and listening are necessary to make conversations. We use Japanese animation which is called “Anime” in the world as a new teaching material of Japanese. Anime has a huge amount of scenes in which characters speak with their voice in the standard and clear pronunciations. It means that Anime has many good examples of Japanese dialogues for JFL learners to improve their speaking and listening. The goal of our research is to propose a new method to classify dialogues in Anime depending on the dialogues’ degrees of difficulty. In this paper, we analyzed the words and expressions used in the Anime dialogues and the scripts of listening tests of the previous Japanese Language Proficiency Tests (JLPT). Each word and expression were classified into different Japanese language levels. We used the analysis results of JLPT as a standard reference to be compared with Anime dialogues. The analysis results showed that the genres of Anime had effects on using words and expressions. We also found that the scripts of listening tests of high-level included many high-level words and expressions.
Junjie Shan, Yoko Nishihara, Ryosuke Yamanishi, Jun-ichi Fukumoto
KES1