Institution profile

Universidade Lusófona de Humanidades e Tecnologias

Academic institutioneurope · pt
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Validity of LLMs as data annotators: AMALIA on authority

Jul 09, 2026

This study investigates whether the Portuguese large language model AMALIA genuinely adheres to theoretical definitions when labeling the moral foundations construct, rather than relying on superficial cues such as “moral outrage.” To this end, the authors introduce a novel metric—recovery gap—that evaluates construct validity by comparing AMALIA-9B’s performance under holistic versus decomposed prompting strategies. Results reveal that while AMALIA performs well with holistic prompts, its performance drops to approximately half under decomposed prompts, indicating a reliance on surface-level correlations. In contrast, open-source multilingual models effectively close this recovery gap. This work represents the first application of the recovery gap methodology to validate moral constructs in a nationally sovereign language model, offering a new paradigm for assessing construct validity in large language models.

0 citationsRead paper

Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

Jun 26, 2026

Although large language models (LLMs) can achieve agreement with human annotators in text coding, their judgments may rely on superficial features unrelated to the underlying theoretical construct, thereby lacking construct validity. To address this issue, this work proposes a “fine-grained calibration” approach that decomposes theoretical constructs into clause-level components, validates each component against extractive evidence, and aggregates results according to explicit theoretical rules to assess whether LLMs genuinely measure the target construct. This method shifts the validation of construct validity from output consistency to process interpretability, enabling identification of errors stemming either from missing components or confusion with neighboring constructs. It establishes a transparent and interpretable paradigm for trustworthy measurement using LLMs in the social sciences.

0 citationsRead paper

Synthetic Data Generation for Screen Time and App Usage

Sep 17, 2025

Collecting real-world smartphone usage data faces significant challenges—including high acquisition costs, privacy risks, limited sample sizes, and severe sampling bias. Method: This paper proposes a large language model (LLM)-based synthetic data generation framework. It introduces a novel three-component prompting schema integrating user persona descriptions, target behavioral constraints, and seed-data guidance. We systematically investigate how prompt granularity and in-context examples affect generation quality. Using models such as ChatGPT, we perform iterative structured data synthesis and rigorously evaluate behavioral plausibility via human assessment and statistical testing. Contribution/Results: Fine-grained prompting substantially improves data reasonableness but entails trade-offs between diversity and fidelity. Our framework establishes a reproducible, evaluable, and scenario-adaptable methodology for generating synthetic behavioral data—enabling scalable, privacy-preserving smartphone usage research without compromising validity.

0 citationsRead paper

HumAIne-Chatbot: Real-Time Personalized Conversational AI via Reinforcement Learning

Sep 04, 2025

Contemporary conversational AI systems lack explicit modeling of individual user differences, hindering dynamic adaptation of both content and stylistic elements. To address this, we propose a reinforcement learning–based personalized dialogue management framework. First, we leverage GPT to generate diverse synthetic personas for policy pre-training. Second, we construct an online user profile that fuses implicit behavioral signals—including typing speed, dwell time, and affective responses—with explicit feedback, and optimize the dialogue policy in real time via multi-step reinforcement learning. Third, we introduce a fine-grained dynamic policy model enabling dual-dimensional adaptation—both stylistic and content-based. Evaluated across 50 synthetic personas, our system achieves significant improvements in user satisfaction (+28.6%), personalization accuracy (+31.4%), and task completion rate (+24.9%), with large effect sizes and high statistical significance (p < 0.001).

0 citationsRead paper

Closing the Loop: A Systematic Review of Experience-Driven Game Adaptation

May 02, 2025

This paper addresses the structural imbalance in existing adaptive game systems—prioritizing performance optimization over affective responsiveness. Through a systematic review of empirically grounded studies from 2015–2024, guided by the PRISMA framework, it identifies critical implementation bottlenecks in the “experience-driven closed loop” across three stages: player perception, affective modeling, and content adaptation. Key findings reveal a severe gap in modeling transient affective states (e.g., stress, anxiety), dominance of knowledge-driven approaches, and underutilization of multimodal affective sensing. To bridge this gap, the paper proposes a novel real-time affective sensing paradigm integrating facial expression analysis with peripheral interaction data. It further underscores the pivotal role of interpretable affective modeling in enhancing immersion and therapeutic efficacy. The work establishes a theoretically grounded framework and actionable design principles for developing truly player experience–centered adaptive game systems.

0 citationsRead paper
Recent publications

Latest Papers

Validity of LLMs as data annotators: AMALIA on authority

Jul 09, 2026

This study investigates whether the Portuguese large language model AMALIA genuinely adheres to theoretical definitions when labeling the moral foundations construct, rather than relying on superficial cues such as “moral outrage.” To this end, the authors introduce a novel metric—recovery gap—that evaluates construct validity by comparing AMALIA-9B’s performance under holistic versus decomposed prompting strategies. Results reveal that while AMALIA performs well with holistic prompts, its performance drops to approximately half under decomposed prompts, indicating a reliance on surface-level correlations. In contrast, open-source multilingual models effectively close this recovery gap. This work represents the first application of the recovery gap methodology to validate moral constructs in a nationally sovereign language model, offering a new paradigm for assessing construct validity in large language models.

0 citationsRead paper

Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

Jun 26, 2026

Although large language models (LLMs) can achieve agreement with human annotators in text coding, their judgments may rely on superficial features unrelated to the underlying theoretical construct, thereby lacking construct validity. To address this issue, this work proposes a “fine-grained calibration” approach that decomposes theoretical constructs into clause-level components, validates each component against extractive evidence, and aggregates results according to explicit theoretical rules to assess whether LLMs genuinely measure the target construct. This method shifts the validation of construct validity from output consistency to process interpretability, enabling identification of errors stemming either from missing components or confusion with neighboring constructs. It establishes a transparent and interpretable paradigm for trustworthy measurement using LLMs in the social sciences.

0 citationsRead paper

Synthetic Data Generation for Screen Time and App Usage

Sep 17, 2025

Collecting real-world smartphone usage data faces significant challenges—including high acquisition costs, privacy risks, limited sample sizes, and severe sampling bias. Method: This paper proposes a large language model (LLM)-based synthetic data generation framework. It introduces a novel three-component prompting schema integrating user persona descriptions, target behavioral constraints, and seed-data guidance. We systematically investigate how prompt granularity and in-context examples affect generation quality. Using models such as ChatGPT, we perform iterative structured data synthesis and rigorously evaluate behavioral plausibility via human assessment and statistical testing. Contribution/Results: Fine-grained prompting substantially improves data reasonableness but entails trade-offs between diversity and fidelity. Our framework establishes a reproducible, evaluable, and scenario-adaptable methodology for generating synthetic behavioral data—enabling scalable, privacy-preserving smartphone usage research without compromising validity.

0 citationsRead paper

HumAIne-Chatbot: Real-Time Personalized Conversational AI via Reinforcement Learning

Sep 04, 2025

Contemporary conversational AI systems lack explicit modeling of individual user differences, hindering dynamic adaptation of both content and stylistic elements. To address this, we propose a reinforcement learning–based personalized dialogue management framework. First, we leverage GPT to generate diverse synthetic personas for policy pre-training. Second, we construct an online user profile that fuses implicit behavioral signals—including typing speed, dwell time, and affective responses—with explicit feedback, and optimize the dialogue policy in real time via multi-step reinforcement learning. Third, we introduce a fine-grained dynamic policy model enabling dual-dimensional adaptation—both stylistic and content-based. Evaluated across 50 synthetic personas, our system achieves significant improvements in user satisfaction (+28.6%), personalization accuracy (+31.4%), and task completion rate (+24.9%), with large effect sizes and high statistical significance (p < 0.001).

0 citationsRead paper

Closing the Loop: A Systematic Review of Experience-Driven Game Adaptation

May 02, 2025

This paper addresses the structural imbalance in existing adaptive game systems—prioritizing performance optimization over affective responsiveness. Through a systematic review of empirically grounded studies from 2015–2024, guided by the PRISMA framework, it identifies critical implementation bottlenecks in the “experience-driven closed loop” across three stages: player perception, affective modeling, and content adaptation. Key findings reveal a severe gap in modeling transient affective states (e.g., stress, anxiety), dominance of knowledge-driven approaches, and underutilization of multimodal affective sensing. To bridge this gap, the paper proposes a novel real-time affective sensing paradigm integrating facial expression analysis with peripheral interaction data. It further underscores the pivotal role of interpretable affective modeling in enhancing immersion and therapeutic efficacy. The work establishes a theoretically grounded framework and actionable design principles for developing truly player experience–centered adaptive game systems.

0 citationsRead paper