SANGAH LEE / PUBLICATIONS
Publications
How Do Language Models Represent and Use Phonological Information for Allomorph Selection?
Accepted to EMNLP 2026
ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
The 2nd International Conference on Human-AI Interaction and Experience Design (HAXD26)
Do Korean-Adapted LLMs Think in Korean? Analyzing Latent Language and the Preservation of Korean-Specific Knowledge
Language and Information, Vol.29, No.3, pp.229–256
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
Findings of the Association for Computational Linguistics: ACL 2025
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
arXiv
A Short Note on the Structural Priming in LLM: Focusing on Dative Constructions in Korean
Language and Information, Vol.28, No.3, pp.111–142 · In Korean
ManWav: The First Manchu ASR Model
Proceedings of the Third Workshop on NLP Applications to Field Linguistics
Large Language Models Show Human-Like Abstract Thinking Patterns: A Construal-Level Perspective
Proceedings of the Annual Meeting of the Cognitive Science Society
ManNER & ManPOS: Pioneering NLP for Endangered Manchu Language
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
arXiv
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression
arXiv
DaG LLM ver 1.0: Pioneering Instruction-Tuned Language Modeling For Korean NLP
arXiv
Mergen: The First Manchu-Korean Machine Translation Model Trained on Augmented Data
3rd Multilingual Representation Learning (MRL) Workshop
Studies on Clauses in Computational Linguistics Focused on Korean Corpora
Journal of Korean Linguistics, No.107, pp.445–468 · In Korean
Contract Eligibility Verification Enhanced by Keyword and Contextual Embeddings
Journal of KIISE, Vol.49, No.10, pp.848–858 · In Korean
The Korean Morphologically Tight-Fitting Tokenizer for Noisy User-Generated Texts
The 7th Workshop on Noisy User-Generated Text (W-NUT 2021)
Combining Sentiment-Combined Model with Pre-Trained BERT Models for Sentiment Analysis
Journal of KIISE, Vol.48, No.7, pp.815-824 · In Korean
Argument Facet Detection in Online Debates Based on Attention Weights and Clustering with Combined Similarity Matrices
Korean Journal of Linguistics, Vol.46, No.1, pp.107–134
KR-BERT: A Small-Scale Korean-Specific BERT Language Model
Journal of KIISE, Vol.47, No.7, pp.682–692
An Analysis of Linear Argumentation Structure of Korean Debate Texts Using Sequential Modeling and Linguistic Features
Journal of KIISE, Vol.45, No.12, pp.1292–1301 · In Korean
Stance Classification of Online Debate Texts based on Discourse Relations
Language Research, Vol.52, No.3, pp.511–532 · In Korean