ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures
研究通过将多代理LLM失败案例转换为功能性伙伴状态推理项目,解决代理间角色、知识或意图跟踪不正确的问题。
研究通过将多代理LLM失败案例转换为功能性伙伴状态推理项目,解决代理间角色、知识或意图跟踪不正确的问题。
为解决儿童视力问题发现晚的问题,本文提出SightSentinel架构,通过教室显示屏定期进行视觉微筛查,并使用多智能体处理数据,以实现早期预警。
Existing AI-based wildfire monitoring systems are prone to false alarms and trust deficits due to the absence of adaptive multi-agent coordination, structured human oversight, and verifiable accountability mechanisms. This work proposes a novel multi-agent architecture integrating blockchain-based governance constraints, formulating the monitoring task as a constrained partially observable Markov decision process (POMDP) with human authorization invariants. The system employs hierarchical coordination to dynamically schedule drones and enforces authorization policies via smart contracts on a permissioned blockchain. Crucially, human authorization is embedded directly into the agents’ decision loop as a state-transition invariant, enabling verifiable accountability. Experimental results demonstrate that the proposed framework maintains high detection performance while significantly reducing false alarm rates and exhibits robustness against injection, replay, and tampering attacks, all with only modest computational overhead.
To address the low fine-tuning efficiency and high GPU memory overhead of LLaMA-3.2-3B on medical chain-of-thought reasoning tasks under GPU and memory resource constraints, this paper proposes a two-stage parameter-efficient fine-tuning method integrating LoRA and QLoRA. By synergistically combining low-rank adaptation with 4-bit quantization, the approach enables lightweight adaptation of the full-parameter model on a single consumer-grade GPU (e.g., RTX 4090). Evaluated on standard medical reasoning benchmarks—MedQA-USMLE and PubMedQA—the method reduces GPU memory consumption by 60% and accelerates training by 2.3× compared to full-parameter fine-tuning, while preserving chain-of-thought coherence and medical factual accuracy (accuracy degradation <1.2%). This significantly enhances the practical deployability of large language models in resource-constrained clinical and medical AI settings.
研究通过将多代理LLM失败案例转换为功能性伙伴状态推理项目,解决代理间角色、知识或意图跟踪不正确的问题。
为解决儿童视力问题发现晚的问题,本文提出SightSentinel架构,通过教室显示屏定期进行视觉微筛查,并使用多智能体处理数据,以实现早期预警。
Existing AI-based wildfire monitoring systems are prone to false alarms and trust deficits due to the absence of adaptive multi-agent coordination, structured human oversight, and verifiable accountability mechanisms. This work proposes a novel multi-agent architecture integrating blockchain-based governance constraints, formulating the monitoring task as a constrained partially observable Markov decision process (POMDP) with human authorization invariants. The system employs hierarchical coordination to dynamically schedule drones and enforces authorization policies via smart contracts on a permissioned blockchain. Crucially, human authorization is embedded directly into the agents’ decision loop as a state-transition invariant, enabling verifiable accountability. Experimental results demonstrate that the proposed framework maintains high detection performance while significantly reducing false alarm rates and exhibits robustness against injection, replay, and tampering attacks, all with only modest computational overhead.
To address the low fine-tuning efficiency and high GPU memory overhead of LLaMA-3.2-3B on medical chain-of-thought reasoning tasks under GPU and memory resource constraints, this paper proposes a two-stage parameter-efficient fine-tuning method integrating LoRA and QLoRA. By synergistically combining low-rank adaptation with 4-bit quantization, the approach enables lightweight adaptation of the full-parameter model on a single consumer-grade GPU (e.g., RTX 4090). Evaluated on standard medical reasoning benchmarks—MedQA-USMLE and PubMedQA—the method reduces GPU memory consumption by 60% and accelerates training by 2.3× compared to full-parameter fine-tuning, while preserving chain-of-thought coherence and medical factual accuracy (accuracy degradation <1.2%). This significantly enhances the practical deployability of large language models in resource-constrained clinical and medical AI settings.