🤖 AI Summary
This work proposes a vision-language model grounded in structured latent variables—specifically Miller indices—to enable physically consistent geometric reasoning about fracture surfaces. By incorporating Miller indices as a latent prior into a multimodal large language model, the method generates crystallographic plane hypotheses through joint visual–language reasoning and validates them against physical constraints. The model not only reliably infers ideal cleavage planes from images but also actively identifies and rejects non-ideal fracture cases that cannot be meaningfully represented by Miller indices. Experimental validation on synthetic data, 2D–3D geometric correspondences, and real-world fracture images across diverse materials—including ceramics, glass, metals, and concrete—demonstrates the approach’s effectiveness and robustness.
📝 Abstract
We study whether multimodal large language models (MLLMs) can leverage crystallographic plane indices (Miller indices) as a structured latent representation for reasoning about fracture geometry. We formulate Miller indices $z = (h,k,l)$ as a latent variable governing idealized planar fracture and evaluate two complementary capabilities: (i) latent inference, where the model maps visual observations to plane hypotheses under physically valid conditions, and (ii) latent applicability assessment, where the model determines whether such a representation is meaningful for a given fracture image.
Through extensive experiments spanning synthetic data, controlled 2D--3D geometric pairs, and real-world fracture images across multiple material classes -- including ceramics, glass, metals, and concrete -- we show that MLLMs can reliably perform latent inference in idealized settings and, critically, can reject the latent representation when the underlying physics does not support it. These results suggest that MLLMs can act as physics-aware reasoning systems conditioned on structured latent priors, provided that the domain of validity is explicitly modeled.