Institution profile

University of Portsmouth

Academic institutioneurope · gb
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

Second-Order Response Laws for LLM Judges: Debiased Estimation of Prompt Instability

Aug 17, 2026

This study addresses estimation bias in large language model evaluation arising from the confounding of prompt instability and sampling noise. We propose a second-order response law framework that derives an unbiased estimator and employs a crossed prompt-answer order experimental design to effectively disentangle inter-prompt variability from intra-prompt sampling noise. This approach enables precise, debiased assessment of prompt robustness under low repetition budgets, yielding corrected estimates that closely approximate high-repetition benchmarks. By resolving noise interference challenges under limited inference calls, this work establishes a novel theoretical foundation and methodological paradigm for cost-effective and reliable LLM evaluation.

0 citationsRead paper

Reveal, Correct, Then Pay: Encrypted Mempools and Perpetual Funding Security

Jul 15, 2026

This work addresses the vulnerability of encrypted mempools to economically lagging and security risks arising from self-authorized state manipulation—such as perpetual contract funding rate manipulation—due to their inability to inject corrective transactions into already committed batches, despite offering protection against victim-dependent MEV attacks. The paper proposes a micro-correction mechanism grounded in executable arbitrage, modeling how correctors optimally choose order sizes balancing price impact and inventory costs, while evaluating exploitable opportunities through the lens of protocol disclosure timing. It introduces a novel local security index incorporating attacker blind spots, correction shielding, and capitalization shielding, revealing how private transactions suppress predictive capitalization of funding rates and induce dual amplification effects. By integrating game theory, market mechanism design, and encrypted mempool architecture, the study establishes a dynamic security framework driven by information scheduling and response factors, proving that closed-phase correction rates fall below adaptive correction rates and quantifying both state distortion and its responsive amplification.

0 citationsRead paper

Slack and Budget Breaking in Threshold Team Production

Jul 07, 2026

This study addresses the incentive challenge in threshold team production, where task failure may result from members delaying submission of their shares. The authors propose a non-negative completion bonus mechanism that relies solely on submitted shares and achieves strong delay immunity without requiring deposits or penalties. By employing a uniform allocation rule and leveraging mechanism design and game-theoretic analysis, the work precisely characterizes the minimal budget required in the worst case under the k-of-n threshold task model. The paper establishes necessary and sufficient conditions to guarantee timely participation by all agents while being robust against collusive delays, and proves that the derived budget bound is tight for all transfer rules based solely on completed shares.

0 citationsRead paper
Recent publications

Latest Papers

Second-Order Response Laws for LLM Judges: Debiased Estimation of Prompt Instability

Aug 17, 2026

This study addresses estimation bias in large language model evaluation arising from the confounding of prompt instability and sampling noise. We propose a second-order response law framework that derives an unbiased estimator and employs a crossed prompt-answer order experimental design to effectively disentangle inter-prompt variability from intra-prompt sampling noise. This approach enables precise, debiased assessment of prompt robustness under low repetition budgets, yielding corrected estimates that closely approximate high-repetition benchmarks. By resolving noise interference challenges under limited inference calls, this work establishes a novel theoretical foundation and methodological paradigm for cost-effective and reliable LLM evaluation.

0 citationsRead paper

Reveal, Correct, Then Pay: Encrypted Mempools and Perpetual Funding Security

Jul 15, 2026

This work addresses the vulnerability of encrypted mempools to economically lagging and security risks arising from self-authorized state manipulation—such as perpetual contract funding rate manipulation—due to their inability to inject corrective transactions into already committed batches, despite offering protection against victim-dependent MEV attacks. The paper proposes a micro-correction mechanism grounded in executable arbitrage, modeling how correctors optimally choose order sizes balancing price impact and inventory costs, while evaluating exploitable opportunities through the lens of protocol disclosure timing. It introduces a novel local security index incorporating attacker blind spots, correction shielding, and capitalization shielding, revealing how private transactions suppress predictive capitalization of funding rates and induce dual amplification effects. By integrating game theory, market mechanism design, and encrypted mempool architecture, the study establishes a dynamic security framework driven by information scheduling and response factors, proving that closed-phase correction rates fall below adaptive correction rates and quantifying both state distortion and its responsive amplification.

0 citationsRead paper

Slack and Budget Breaking in Threshold Team Production

Jul 07, 2026

This study addresses the incentive challenge in threshold team production, where task failure may result from members delaying submission of their shares. The authors propose a non-negative completion bonus mechanism that relies solely on submitted shares and achieves strong delay immunity without requiring deposits or penalties. By employing a uniform allocation rule and leveraging mechanism design and game-theoretic analysis, the work precisely characterizes the minimal budget required in the worst case under the k-of-n threshold task model. The paper establishes necessary and sufficient conditions to guarantee timely participation by all agents while being robust against collusive delays, and proves that the derived budget bound is tight for all transfer rules based solely on completed shares.

0 citationsRead paper