Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
本文提出一种高效管道,通过去除冗余、语言识别和质量分类等方法,从葡萄牙语网页中筛选出高质量的欧洲葡萄牙语文本数据集。
本文提出一种高效管道,通过去除冗余、语言识别和质量分类等方法,从葡萄牙语网页中筛选出高质量的欧洲葡萄牙语文本数据集。
This study reveals a critical blind spot in current large vision-language models (LVLMs): their inability to reliably assess the temporal and logical coherence of image sequences, particularly in detecting semantically misordered or contradictory narratives. Through systematic evaluation on pairwise temporal discrimination tasks, we uncover a structural failure in LVLMs driven by primacy and recency effects, causing them to over-rely on frame position rather than genuine narrative logic. This bias stems from architectural choices such as causal masking and rotary embeddings. Diagnostic probes, positional perturbations, and temporal discrimination experiments demonstrate that while LVLMs perform adequately on frame-level scoring, their performance sharply degrades in long-range temporal reasoning. These findings challenge the prevailing snapshot-centric evaluation paradigm and question the suitability of LVLMs as trustworthy evaluators of visual narrative coherence.
This study investigates the computational complexity and formal language-theoretic properties of the language of shortest representatives of conjugacy classes in free inverse monoids. It extends the notion of conjugacy languages to this algebraic setting by introducing a new equivalence relation, UConj, which distinguishes between trivial and nontrivial conjugacy classes. Employing techniques from formal language theory, automata theory, and combinatorial group theory, the authors prove that for rank at least two, this language is neither context-free nor co-context-free, whereas in the one-generator case it is context-free. Furthermore, they provide an explicit context-free grammar for the geodesic language of trivial conjugacy classes. These results are subsequently generalized to hyperbolic groups, right-angled Artin groups, and virtually abelian groups.
This work proposes a unified implicit computational framework based on discrete ordinary differential equations (ODEs) to characterize the essential features of complexity classes such as counting classes (e.g., ⊕P) and alternating classes (e.g., the polynomial hierarchy). Starting from a base class weaker than FP, the framework employs a single underlying algebra and three fundamental recursion schemes, using the nesting depth of ODE operators to precisely capture complexity hierarchies. It extends discrete ODE methods—previously confined to deterministic and nondeterministic settings—to the realm of counting complexity for the first time, revealing the computational content under linear constraints and establishing a theoretical bridge between differential mechanisms and counting or alternation processes. This approach offers a novel paradigm for implicit complexity theory, enabling a unified representation of diverse complexity classes and facilitating its extension to broader classes.
This study addresses the critical absence of European Portuguese data in existing OCR benchmarks tailored to modern application scenarios. To bridge this gap, the authors introduce PorTEXTO, the first benchmark specifically designed for contemporary cultural contexts in European Portuguese, comprising real-world images paired with high-quality text annotations. These annotations are produced through a hybrid pipeline combining automatic generation by large vision-language models with manual verification by native speakers. The work fills a significant void in authentic-scenario OCR evaluation for this language and releases all resources publicly. Experimental results demonstrate that state-of-the-art OCR models suffer substantial performance degradation on real-world samples, and that incorporating dedicated multilingual training data yields greater improvements in recognition accuracy than merely increasing model scale or input resolution.
本文提出一种高效管道,通过去除冗余、语言识别和质量分类等方法,从葡萄牙语网页中筛选出高质量的欧洲葡萄牙语文本数据集。
This study reveals a critical blind spot in current large vision-language models (LVLMs): their inability to reliably assess the temporal and logical coherence of image sequences, particularly in detecting semantically misordered or contradictory narratives. Through systematic evaluation on pairwise temporal discrimination tasks, we uncover a structural failure in LVLMs driven by primacy and recency effects, causing them to over-rely on frame position rather than genuine narrative logic. This bias stems from architectural choices such as causal masking and rotary embeddings. Diagnostic probes, positional perturbations, and temporal discrimination experiments demonstrate that while LVLMs perform adequately on frame-level scoring, their performance sharply degrades in long-range temporal reasoning. These findings challenge the prevailing snapshot-centric evaluation paradigm and question the suitability of LVLMs as trustworthy evaluators of visual narrative coherence.
This study investigates the computational complexity and formal language-theoretic properties of the language of shortest representatives of conjugacy classes in free inverse monoids. It extends the notion of conjugacy languages to this algebraic setting by introducing a new equivalence relation, UConj, which distinguishes between trivial and nontrivial conjugacy classes. Employing techniques from formal language theory, automata theory, and combinatorial group theory, the authors prove that for rank at least two, this language is neither context-free nor co-context-free, whereas in the one-generator case it is context-free. Furthermore, they provide an explicit context-free grammar for the geodesic language of trivial conjugacy classes. These results are subsequently generalized to hyperbolic groups, right-angled Artin groups, and virtually abelian groups.
This work proposes a unified implicit computational framework based on discrete ordinary differential equations (ODEs) to characterize the essential features of complexity classes such as counting classes (e.g., ⊕P) and alternating classes (e.g., the polynomial hierarchy). Starting from a base class weaker than FP, the framework employs a single underlying algebra and three fundamental recursion schemes, using the nesting depth of ODE operators to precisely capture complexity hierarchies. It extends discrete ODE methods—previously confined to deterministic and nondeterministic settings—to the realm of counting complexity for the first time, revealing the computational content under linear constraints and establishing a theoretical bridge between differential mechanisms and counting or alternation processes. This approach offers a novel paradigm for implicit complexity theory, enabling a unified representation of diverse complexity classes and facilitating its extension to broader classes.
This study addresses the critical absence of European Portuguese data in existing OCR benchmarks tailored to modern application scenarios. To bridge this gap, the authors introduce PorTEXTO, the first benchmark specifically designed for contemporary cultural contexts in European Portuguese, comprising real-world images paired with high-quality text annotations. These annotations are produced through a hybrid pipeline combining automatic generation by large vision-language models with manual verification by native speakers. The work fills a significant void in authentic-scenario OCR evaluation for this language and releases all resources publicly. Experimental results demonstrate that state-of-the-art OCR models suffer substantial performance degradation on real-world samples, and that incorporating dedicated multilingual training data yields greater improvements in recognition accuracy than merely increasing model scale or input resolution.