Institution profile

Semyung University

Academic institutionasia · kr
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams

Mar 28, 2026

This work addresses the absence of evaluation benchmarks for vision-language models in real-world multilingual, multimodal civil service examination scenarios. The authors introduce a novel benchmark derived from authentic civil service exams across five Eurasian regions, comprising over 8,000 high-resolution scanned questions. For the first time, the benchmark uses full-page exam document images as input, integrating multilingual text, tables, and complex layouts while emphasizing cultural authenticity and visual complexity. High-fidelity OCR preserves original document structure, and a standardized instruction format enables joint vision–language modeling and layout-aware reasoning. Experimental results reveal that even state-of-the-art models achieve only 86% accuracy, underscoring the benchmark’s difficulty and its value for advancing e-governance, public document analysis, and equitable test preparation.

0 citationsRead paper

How Language Directions Align with Token Geometry in Multilingual LLMs

Nov 16, 2025

This study investigates how linguistic information is structurally encoded and evolves across layers in multilingual large language models (LLMs). To this end, we propose the Token-Language Alignment (TLA) analytical framework, which employs both linear and nonlinear probing techniques to perform layer-wise dynamic modeling across all 268 Transformer layers. We find that language identity is highly separable as early as the first layer (accuracy: 76.4±8.2%), and remains approximately linearly separable throughout the entire depth. Crucially, we uncover for the first time that the alignment between language-specific directional representations and vocabulary embeddings is significantly modulated by the language composition of the training corpus—particularly, Chinese-dominant models exhibit a strong “structural imprint effect.” Our work establishes an interpretable, quantifiable analytical paradigm for characterizing the internal representational mechanisms of multilingual LLMs, enabling systematic investigation of cross-lingual representation geometry and its dependence on pretraining data distribution.

0 citationsRead paper

KoTaP: A Panel Dataset for Corporate Tax Avoidance, Performance, and Governance in Korea

Nov 06, 2025

This paper addresses the scarcity of high-quality, localized data for corporate tax avoidance research in Korea. To this end, we construct KoTaP—the first long-term, standardized, balanced panel dataset for Korea (2011–2024)—covering non-financial firms listed on KOSPI/KOSDAQ, comprising 12,653 firm-year observations. KoTaP systematically integrates multidimensional variables: tax avoidance measures (cash- and GAAP-based effective tax rates, book-tax differences), earnings management, profitability (ROA/ROE), leverage (LEV), firm size (SIZE), audit quality (BIG4), and corporate governance, while explicitly incorporating Korean institutional features—including concentrated ownership and high foreign ownership. Designed to balance international comparability with local contextual fidelity, KoTaP supports rigorous econometric modeling, deep learning applications, and interpretable AI analysis. The dataset is publicly available, serving as critical infrastructure for empirical accounting and finance research, external validity testing, audit practice optimization, and evidence-based investment and policy decisions.

0 citationsRead paper
Recent publications

Latest Papers

EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams

Mar 28, 2026

This work addresses the absence of evaluation benchmarks for vision-language models in real-world multilingual, multimodal civil service examination scenarios. The authors introduce a novel benchmark derived from authentic civil service exams across five Eurasian regions, comprising over 8,000 high-resolution scanned questions. For the first time, the benchmark uses full-page exam document images as input, integrating multilingual text, tables, and complex layouts while emphasizing cultural authenticity and visual complexity. High-fidelity OCR preserves original document structure, and a standardized instruction format enables joint vision–language modeling and layout-aware reasoning. Experimental results reveal that even state-of-the-art models achieve only 86% accuracy, underscoring the benchmark’s difficulty and its value for advancing e-governance, public document analysis, and equitable test preparation.

0 citationsRead paper

How Language Directions Align with Token Geometry in Multilingual LLMs

Nov 16, 2025

This study investigates how linguistic information is structurally encoded and evolves across layers in multilingual large language models (LLMs). To this end, we propose the Token-Language Alignment (TLA) analytical framework, which employs both linear and nonlinear probing techniques to perform layer-wise dynamic modeling across all 268 Transformer layers. We find that language identity is highly separable as early as the first layer (accuracy: 76.4±8.2%), and remains approximately linearly separable throughout the entire depth. Crucially, we uncover for the first time that the alignment between language-specific directional representations and vocabulary embeddings is significantly modulated by the language composition of the training corpus—particularly, Chinese-dominant models exhibit a strong “structural imprint effect.” Our work establishes an interpretable, quantifiable analytical paradigm for characterizing the internal representational mechanisms of multilingual LLMs, enabling systematic investigation of cross-lingual representation geometry and its dependence on pretraining data distribution.

0 citationsRead paper

KoTaP: A Panel Dataset for Corporate Tax Avoidance, Performance, and Governance in Korea

Nov 06, 2025

This paper addresses the scarcity of high-quality, localized data for corporate tax avoidance research in Korea. To this end, we construct KoTaP—the first long-term, standardized, balanced panel dataset for Korea (2011–2024)—covering non-financial firms listed on KOSPI/KOSDAQ, comprising 12,653 firm-year observations. KoTaP systematically integrates multidimensional variables: tax avoidance measures (cash- and GAAP-based effective tax rates, book-tax differences), earnings management, profitability (ROA/ROE), leverage (LEV), firm size (SIZE), audit quality (BIG4), and corporate governance, while explicitly incorporating Korean institutional features—including concentrated ownership and high foreign ownership. Designed to balance international comparability with local contextual fidelity, KoTaP supports rigorous econometric modeling, deep learning applications, and interpretable AI analysis. The dataset is publicly available, serving as critical infrastructure for empirical accounting and finance research, external validity testing, audit practice optimization, and evidence-based investment and policy decisions.

0 citationsRead paper