CBIL: Collective Behavior Imitation Learning for Fish from Real Videos
This work addresses the challenge of modeling collective behaviors in high-density, irregular fish schools. We propose the first end-to-end, video-driven self-supervised imitation learning framework that learns spatiotemporal motion patterns directly from raw videos—without requiring ground-truth trajectory annotations. Methodologically, the framework integrates a Masked Video Autoencoder (MVAE) with self-supervised video representation learning to extract robust spatiotemporal features; introduces an adversarial latent-space motion distribution matching mechanism; and incorporates biologically inspired reward functions and prior motion constraints to enhance training stability and behavioral plausibility. Experiments demonstrate substantial improvements in behavioral diversity and visual fidelity of generated motions. The framework generalizes effectively to multi-species animated synthesis and enables automatic detection of anomalous schooling behaviors in field-captured videos.