Institution profile

University of the Faroe Islands

Academic institutioneurope · fo
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal

Sep 01, 2025

This work addresses the lack of controllable refusal capability in large language models (LLMs). We propose the first unified framework for uncertainty calibration and risk-controlled rejection tailored to API-based black-box LLMs—requiring no fine-tuning and offering distribution-free theoretical guarantees. Methodologically, it integrates heterogeneous uncertainty signals—including sequence likelihood, self-consistency dispersion, retrieval compatibility, and tool feedback—into a lightweight calibration via temperature scaling and adaptive scoring, then enforces principled rejection using conformal risk control under user-specified error budgets. Key innovations include fine-grained factual alignment and interpretable refusal. Experiments across short-form QA, code generation, and retrieval-augmented long-text generation demonstrate substantial improvements over entropy- and logit-threshold baselines: lower calibration error, superior area under the risk–coverage curve, and higher coverage at fixed risk levels.

0 citationsRead paper

The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases

Aug 17, 2025

Large language models (LLMs) implicitly inherit cultural values from their training corpora, yet systematic, quantifiable assessment of such value alignment remains lacking. Method: We introduce the concept of “cultural memes”—systematically internalized value orientations—and focus on two cross-cultural dimensions: Individualism–Collectivism (IDV) and Power Distance (PDI). We construct the Cultural Probe Dataset (CPD) and propose the Cultural Alignment Index (CAI), evaluating LLMs via zero-shot prompting, human annotation, statistical testing, and correlation with Hofstede’s national culture scores. Contribution/Results: Experiments reveal GPT-4 exhibits strong individualist and low-power-distance preferences, whereas ERNIE Bot favors collectivism and high power distance; both achieve CAI > 0.8 (p < 0.001), confirming LLMs statistically mirror the cultural profiles of their training data. This work establishes the first measurable, comparable, and interpretable framework for assessing cultural value alignment in LLMs.

0 citationsRead paper
Recent publications

Latest Papers

Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal

Sep 01, 2025

This work addresses the lack of controllable refusal capability in large language models (LLMs). We propose the first unified framework for uncertainty calibration and risk-controlled rejection tailored to API-based black-box LLMs—requiring no fine-tuning and offering distribution-free theoretical guarantees. Methodologically, it integrates heterogeneous uncertainty signals—including sequence likelihood, self-consistency dispersion, retrieval compatibility, and tool feedback—into a lightweight calibration via temperature scaling and adaptive scoring, then enforces principled rejection using conformal risk control under user-specified error budgets. Key innovations include fine-grained factual alignment and interpretable refusal. Experiments across short-form QA, code generation, and retrieval-augmented long-text generation demonstrate substantial improvements over entropy- and logit-threshold baselines: lower calibration error, superior area under the risk–coverage curve, and higher coverage at fixed risk levels.

0 citationsRead paper

The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases

Aug 17, 2025

Large language models (LLMs) implicitly inherit cultural values from their training corpora, yet systematic, quantifiable assessment of such value alignment remains lacking. Method: We introduce the concept of “cultural memes”—systematically internalized value orientations—and focus on two cross-cultural dimensions: Individualism–Collectivism (IDV) and Power Distance (PDI). We construct the Cultural Probe Dataset (CPD) and propose the Cultural Alignment Index (CAI), evaluating LLMs via zero-shot prompting, human annotation, statistical testing, and correlation with Hofstede’s national culture scores. Contribution/Results: Experiments reveal GPT-4 exhibits strong individualist and low-power-distance preferences, whereas ERNIE Bot favors collectivism and high power distance; both achieve CAI > 0.8 (p < 0.001), confirming LLMs statistically mirror the cultural profiles of their training data. This work establishes the first measurable, comparable, and interpretable framework for assessing cultural value alignment in LLMs.

0 citationsRead paper