WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition
为解决野外视频数据标注成本高及现有模型在复杂海洋环境中的失效问题,本文通过创建WildFin数据集并使用计算机视觉方法进行鱼类行为识别。
为解决野外视频数据标注成本高及现有模型在复杂海洋环境中的失效问题,本文通过创建WildFin数据集并使用计算机视觉方法进行鱼类行为识别。
This work proposes an agent-centric autoregressive generative model to algorithmically elucidate the mechanisms of perception, internal modeling, planning, and action generation in individual and collective animal behavior. Operating in an egocentric coordinate frame, the model processes pose trajectory data and predicts discretized action sequences to naturally emulate the behavioral constraints inherent to an organism’s self-referential perspective. The framework supports parallel multi-agent representations, enabling complex social behaviors to emerge spontaneously from each agent’s independent perception of and response to others. An accompanying open-source library provides composable action representations and systematic evaluation tools. Experiments demonstrate that the model accurately captures the distribution of social behaviors in courting fruit fly groups and exhibits strong cross-domain transferability.
Current AI systems struggle to autonomously execute end-to-end scientific discovery pipelines in neuroscience, particularly due to a lack of capability in assessing scientific plausibility. This work presents the first systematic evaluation of general-purpose code-generating agents on a real-world, large-scale optogenetics task in Drosophila, whose scale and complexity substantially exceed existing benchmarks. By integrating iterative code generation, visualization of intermediate outputs, and rigorous domain-expert-driven assessment, the study reveals that while agents can successfully complete individual pipeline stages, they fail to reliably orchestrate the full workflow. Performance degrades markedly in the absence of predefined evaluation criteria. These findings highlight critical limitations in AI’s capacity for scientific reasoning and self-evaluation, and propose principles for constructing and evaluating agent-based approaches to open-ended scientific problems.
This study addresses the challenge of simultaneously extracting abstract relational structures from continuous high-dimensional dynamic experiences and enabling cross-context knowledge transfer. Inspired by the neural mechanisms of the hippocampus (HPC) and medial entorhinal cortex (MEC), the work proposes the first self-supervised hierarchical world model that functionally disentangles these two brain regions: an inverse model captures relational structure, while velocity-driven path integration decouples structural representations in the MEC from episodic details encoded in the HPC. This framework unifies structure discovery and reuse, demonstrating superior abstraction capabilities on raw transformation-based dynamic tasks and achieving robust prediction and generalization across multiple contexts.
Neuroscience data are highly fragmented due to heterogeneous experimental paradigms and formats, severely impeding reuse. This study presents the first systematic evaluation of large language model–based AI agents in end-to-end reuse of real-world neural data. By integrating scientific papers, code, and datasets through prompt engineering, the agent automatically parses and reformats multi-source data to support neural–behavioral decoding tasks. The results show that while the agent can successfully execute individual subtasks, it struggles to complete the entire pipeline without errors. Moreover, its reliability as an evaluator is limited in the absence of ground truth. These findings highlight critical challenges for neural data sharing in the AI era and propose best practices centered on human–AI collaboration.
为解决野外视频数据标注成本高及现有模型在复杂海洋环境中的失效问题,本文通过创建WildFin数据集并使用计算机视觉方法进行鱼类行为识别。
This work proposes an agent-centric autoregressive generative model to algorithmically elucidate the mechanisms of perception, internal modeling, planning, and action generation in individual and collective animal behavior. Operating in an egocentric coordinate frame, the model processes pose trajectory data and predicts discretized action sequences to naturally emulate the behavioral constraints inherent to an organism’s self-referential perspective. The framework supports parallel multi-agent representations, enabling complex social behaviors to emerge spontaneously from each agent’s independent perception of and response to others. An accompanying open-source library provides composable action representations and systematic evaluation tools. Experiments demonstrate that the model accurately captures the distribution of social behaviors in courting fruit fly groups and exhibits strong cross-domain transferability.
Current AI systems struggle to autonomously execute end-to-end scientific discovery pipelines in neuroscience, particularly due to a lack of capability in assessing scientific plausibility. This work presents the first systematic evaluation of general-purpose code-generating agents on a real-world, large-scale optogenetics task in Drosophila, whose scale and complexity substantially exceed existing benchmarks. By integrating iterative code generation, visualization of intermediate outputs, and rigorous domain-expert-driven assessment, the study reveals that while agents can successfully complete individual pipeline stages, they fail to reliably orchestrate the full workflow. Performance degrades markedly in the absence of predefined evaluation criteria. These findings highlight critical limitations in AI’s capacity for scientific reasoning and self-evaluation, and propose principles for constructing and evaluating agent-based approaches to open-ended scientific problems.
This study addresses the challenge of simultaneously extracting abstract relational structures from continuous high-dimensional dynamic experiences and enabling cross-context knowledge transfer. Inspired by the neural mechanisms of the hippocampus (HPC) and medial entorhinal cortex (MEC), the work proposes the first self-supervised hierarchical world model that functionally disentangles these two brain regions: an inverse model captures relational structure, while velocity-driven path integration decouples structural representations in the MEC from episodic details encoded in the HPC. This framework unifies structure discovery and reuse, demonstrating superior abstraction capabilities on raw transformation-based dynamic tasks and achieving robust prediction and generalization across multiple contexts.
Neuroscience data are highly fragmented due to heterogeneous experimental paradigms and formats, severely impeding reuse. This study presents the first systematic evaluation of large language model–based AI agents in end-to-end reuse of real-world neural data. By integrating scientific papers, code, and datasets through prompt engineering, the agent automatically parses and reformats multi-source data to support neural–behavioral decoding tasks. The results show that while the agent can successfully execute individual subtasks, it struggles to complete the entire pipeline without errors. Moreover, its reliability as an evaluator is limited in the absence of ground truth. These findings highlight critical challenges for neural data sharing in the AI era and propose best practices centered on human–AI collaboration.