Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu
This study utilizes contextual embeddings from multiple BERT models to provide computational evidence that Urdu light verbs are systematically distinct from their main verb counterparts in event structure while maintaining lemma-specific lexical relatedness, consistent with Butt's theoretical analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Contextual Embedding Evidence for Main–Light Verb Distinctions in Urdu
Problem and Motivation
Urdu grammar features light verb constructions (LVCs) where a lexical verb (V1) combines with a second verb (V2) that contributes aspectual, completive, or event-structural meaning while exhibiting reduced lexical content. While linguistic theory, particularly the work of Butt (1995, 2003, 2010) and Butt and Lahiri (2013), posits that light verbs are distinct from their main verb counterparts yet retain a specific lexical relationship, it remains unclear whether contextual language models encode this distinction. Specifically, it is unknown if the representational geometry of encoder-based models reflects the theoretical predictions that: (1) main and light uses are systematically differentiated, (2) same-lemma uses remain closer to each other than mismatched lemma pairs, and (3) individual light verbs retain distinguishable profiles rather than collapsing into a homogeneous class.
Methodology
The study analyzes 1,126 naturally occurring Urdu sentences containing seven canonical light verbs (āyā, uṭhā, baiṭhā, paṛā, diyā, gayā, liyā). The dataset was constructed from a 1.6 GB in-house corpus, with sentences extracted where the target verb appeared in sentence-final position. Annotation followed a strict three-way agreement protocol: GPT-5.1 first classified each sentence using prompts grounded in Butt's linguistic criteria, and then two expert human annotators independently reviewed the GPT-assigned labels. Only sentences for which all three sources (GPT-5.1, Annotator 1, and Annotator 2) were in full agreement were retained; any sentence with even a single disagreement was discarded.
Three encoder-based models were evaluated:
- UrduBERT: A BERT-base model trained from scratch on a 5.8 GB deduplicated Urdu corpus.
- DunbaaBERT: A RoBERTa-style model trained on a 17 GB Urdu corpus.
- Multilingual BERT (mBERT): A standard multilingual model not specifically optimized for Urdu.
The methodology employed three complementary analyses on contextual embeddings extracted from the final hidden layers:
- Representational Separation: Cosine distance and silhouette scores were calculated between main-use and light-use centroids for each verb and model. Statistical significance was assessed via label-permutation tests.
- Lexical Relatedness: The distance between same-lemma main–light centroids was compared against mismatched cross-lemma distances to test if same-lemma pairs remain closer.
- Verb-Specific Structure: A seven-way classification task was performed to predict light-verb identity. This included a "masked-target" condition where the target verb was replaced with a mask token, and a "preceding-form-disjoint" evaluation to test generalization beyond repeated V1–V2 combinations.
Key Results
- Systematic Differentiation: All 21 verb–model comparisons showed significant representational separation between main and light uses (). UrduBERT demonstrated the largest centroid distances (mean 0.177) and highest silhouette scores (mean 0.272), followed by mBERT and DunbaaBERT.
- Linear and Unsupervised Recoverability: Logistic probing achieved high accuracy in distinguishing main from light uses across all models (UrduBERT mean F1: 0.997). However, unsupervised KMeans clustering showed that while UrduBERT achieved a mean Adjusted Rand Index (ARI) of 0.778, the other models performed near chance, indicating that high linear separability does not guarantee the formation of two dominant spherical clusters.
- Lexical Relatedness: Same-lemma main–light centroids were consistently closer than mismatched cross-lemma pairs across all models. Top-1 matching accuracy was 1.0 for all models, supporting the prediction that main and light uses retain lexical relatedness despite their differentiation.
- Light Verb Identity: In the masked-target condition, UrduBERT achieved 0.866 accuracy and 0.852 macro-F1 in predicting the specific light verb identity, significantly outperforming chance. Under the preceding-form-disjoint evaluation, UrduBERT retained 0.782 accuracy, suggesting generalization beyond immediate surface co-occurrences. DunbaaBERT and mBERT showed lower but still significant performance in masked conditions.
Significance and Claims
The paper provides computational evidence consistent with Butt's linguistic account of Urdu light verbs. The findings demonstrate that contextual embeddings systematically differentiate main and light uses while preserving lemma-specific structure. Specifically, the results support the view that light verbs are not semantically empty grammatical markers but retain verb-specific contributions and a relationship to their main verb counterparts.
The authors explicitly state that these findings provide evidence for representational consequences of Butt's analysis without directly establishing a particular formal lexical-entry structure or diachronic pathway. The study complements previous work on English light verbs by extending the investigation to Urdu and evaluating specific predictions regarding lexical relatedness and verb-specificity that are central to Urdu linguistic theory. The results suggest that encoder-based models, particularly those trained on monolingual Urdu data, capture the nuanced representational geometry of light verb constructions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.