Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints
本文提出集体反事实规划模型,通过协调、同意和验证解决团队在表示几何限制下的共同目标实现问题。
本文提出集体反事实规划模型,通过协调、同意和验证解决团队在表示几何限制下的共同目标实现问题。
This study addresses the low speech recognition accuracy of Burmese medical audio in existing large models, primarily due to data scarcity in clinical settings and environmental noise. To tackle this challenge, we present the first high-quality Burmese medical speech corpus comprising 28 hours of native-speaker-recorded and validated utterances. Leveraging this dataset, we fine-tune the Whisper model using both full-parameter fine-tuning (FFT) and parameter-efficient LoRA-based fine-tuning (PEFT), further enhancing robustness through waveform- and spectrogram-level data augmentation to mitigate the effects of noise and reverberation. Our best-performing system, myMediWhisper-Medium, achieves a word error rate (WER) of 23.44% without data augmentation, significantly outperforming larger general-domain fine-tuned models and establishing a new state-of-the-art result for Burmese medical speech recognition.
This study addresses the limitations of traditional RNNs and LSTMs in modeling continuous-time sequences with sparse or missing data, such as clinical physiological signals. It systematically evaluates liquid neural networks (LNNs), particularly the Closed-form Continuous-time (CfC) model, on multimodal native time-series tasks by modeling hidden-state dynamics through continuous differential equations and incorporating temporal dropout for robustness stress testing. Experiments across diverse datasets—including N-MNIST, QuickDraw, IAM, and PhysioNet Sepsis-3—demonstrate that LNNs significantly outperform LSTMs in both parameter efficiency and resilience to missing data. The results highlight the superior practical utility of LNNs in real-world applications such as event-based vision, handwriting recognition, and clinical monitoring.
This work proposes a compact cognitive architecture that unifies the modeling of an agent’s cognitive limitations in explanation, learning, and empathy through a single scalar residual signal driving explanation, decision-making, and accountable refusal behaviors. The framework integrates three modes of cognitive failure into a residual-sufficiency constraint, combining an explanation-decision unit, a family of local representational schemes, and a description-length-driven expansion strategy to enable honest responses to novel situations. The architecture is theoretically proven to be total, deterministic, and guaranteed to terminate uniquely within finitely many steps. Empirical evaluations reproduce three hallmark phenomena—“unknown-type” responses, “bounded empathy,” and “developmental prerequisites”—and yield falsifiable predictions. Notably, this is the first implementation of a typed, witness-augmented, and accountable abandonment mechanism.
This study addresses the lack of a systematic benchmark for Burmese handwritten digit recognition, which has hindered related AI research. The authors establish the first cross-paradigm benchmark on the myMNIST dataset, comprehensively evaluating eleven models—including CNNs, MLPs, LSTMs, GRUs, Transformers, FastKAN, EfficientKAN, JEM, and PETNN variants with various activation functions. Experimental results demonstrate that CNNs achieve the highest performance (accuracy: 0.9970, F1-score: 0.9959), closely followed by PETNN with GELU activation (accuracy: 0.9966), both significantly outperforming KAN-based models and Transformers; JEM also shows competitive results. This work not only fills a critical gap in Burmese script recognition benchmarks but also highlights the potential of physics-inspired PETNN models for regional script recognition and quantifies the performance gap between energy-inspired and true energy-based models.
本文提出集体反事实规划模型,通过协调、同意和验证解决团队在表示几何限制下的共同目标实现问题。
This study addresses the low speech recognition accuracy of Burmese medical audio in existing large models, primarily due to data scarcity in clinical settings and environmental noise. To tackle this challenge, we present the first high-quality Burmese medical speech corpus comprising 28 hours of native-speaker-recorded and validated utterances. Leveraging this dataset, we fine-tune the Whisper model using both full-parameter fine-tuning (FFT) and parameter-efficient LoRA-based fine-tuning (PEFT), further enhancing robustness through waveform- and spectrogram-level data augmentation to mitigate the effects of noise and reverberation. Our best-performing system, myMediWhisper-Medium, achieves a word error rate (WER) of 23.44% without data augmentation, significantly outperforming larger general-domain fine-tuned models and establishing a new state-of-the-art result for Burmese medical speech recognition.
This study addresses the limitations of traditional RNNs and LSTMs in modeling continuous-time sequences with sparse or missing data, such as clinical physiological signals. It systematically evaluates liquid neural networks (LNNs), particularly the Closed-form Continuous-time (CfC) model, on multimodal native time-series tasks by modeling hidden-state dynamics through continuous differential equations and incorporating temporal dropout for robustness stress testing. Experiments across diverse datasets—including N-MNIST, QuickDraw, IAM, and PhysioNet Sepsis-3—demonstrate that LNNs significantly outperform LSTMs in both parameter efficiency and resilience to missing data. The results highlight the superior practical utility of LNNs in real-world applications such as event-based vision, handwriting recognition, and clinical monitoring.
This work proposes a compact cognitive architecture that unifies the modeling of an agent’s cognitive limitations in explanation, learning, and empathy through a single scalar residual signal driving explanation, decision-making, and accountable refusal behaviors. The framework integrates three modes of cognitive failure into a residual-sufficiency constraint, combining an explanation-decision unit, a family of local representational schemes, and a description-length-driven expansion strategy to enable honest responses to novel situations. The architecture is theoretically proven to be total, deterministic, and guaranteed to terminate uniquely within finitely many steps. Empirical evaluations reproduce three hallmark phenomena—“unknown-type” responses, “bounded empathy,” and “developmental prerequisites”—and yield falsifiable predictions. Notably, this is the first implementation of a typed, witness-augmented, and accountable abandonment mechanism.
This study addresses the lack of a systematic benchmark for Burmese handwritten digit recognition, which has hindered related AI research. The authors establish the first cross-paradigm benchmark on the myMNIST dataset, comprehensively evaluating eleven models—including CNNs, MLPs, LSTMs, GRUs, Transformers, FastKAN, EfficientKAN, JEM, and PETNN variants with various activation functions. Experimental results demonstrate that CNNs achieve the highest performance (accuracy: 0.9970, F1-score: 0.9959), closely followed by PETNN with GELU activation (accuracy: 0.9966), both significantly outperforming KAN-based models and Transformers; JEM also shows competitive results. This work not only fills a critical gap in Burmese script recognition benchmarks but also highlights the potential of physics-inspired PETNN models for regional script recognition and quantifies the performance gap between energy-inspired and true energy-based models.