Stereotypical gender actions can be extracted from web text
This study investigates whether web-based text—particularly Twitter corpora—can effectively represent gender-stereotyped behavioral associations and align with human commonsense judgments. Methodologically, we propose a quantification framework for action-gender bias grounded in user gender metadata and pronoun/name heuristics, integrated with the Open Mind Common Sense knowledge base to systematically extract and annotate gender-associated actions. To our knowledge, this is the first cross-source validation of gendered behavioral stereotypes between web text and structured commonsense knowledge. We construct a high-quality gender-action dataset comprising 441 manually annotated and 21,442 automatically annotated instances. Experimental results show that our model achieves a Spearman correlation of 0.47 and an AUC of 0.76 against human-annotated gold standards, demonstrating that web text can robustly model and complement commonsense-level gendered behavioral stereotypes with high recall. This work establishes a novel paradigm for large-scale, dynamic modeling of gender cognition.