Article In: International Journal of Corpus Linguistics: Online-First Articles
A large-scale pipeline for LLM-assisted corpus annotation
Variation and change in the English consider construction
This content is being prepared for publication; it may be subject to changes.
Abstract
As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical annotation using large language models (LLMs). We demonstrate the pipeline’s accessibility and effectiveness through a diachronic case study of variation in the English evaluative consider construction (consider X as/to be/Ø Y). We annotate 143,933 ‘consider’ concordance lines from the Corpus of Historical American English (COHA) via the OpenAI API in under 60 hours, achieving 98%+ accuracy on two sophisticated annotation procedures. Subsequent analysis reveals previously undocumented genre-specific trajectories of change, enabling us to advance new hypotheses about the relationship between register formality and competing pressures of morphosyntactic reduction and enhancement. Our results suggest that LLMs can perform a range of data preparation tasks at scale with minimal human intervention, unlocking substantive research questions previously beyond practical reach.
Article outline
- 1.Introduction
- 2.LLM-assisted corpus annotation and the consider construction
- 2.1LLM-assisted corpus annotation
- 2.2A case study of complementizer variation in evaluative consider constructions
- 2.3An iterative, supervised training pipeline for LLM-based corpus annotation
- 3.The annotation pipeline and its implementation in COHA
- 3.1A four-phase pipeline for automatic corpus annotation
- 3.2Case study implementation: Evaluative consider in COHA
- 3.2.1Data
- 3.2.2LLM selection and implementation
- 3.2.3Task-specific prompt engineering
- 3.2.4API configuration details
- 4.Results: Pipeline performance and diachronic findings
- 4.1Classification accuracy and computational efficiency
- 4.2Error analysis
- 4.3Linguistic findings
- 5.Discussion and conclusion
- Notes
References
References (52)
Anthropic. (2025, September 29). Introducing Claude Sonnet 4.5. Anthropic.com. [URL]
Atıl, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., & Baldwin, B. (2025). Non-determinism of “deterministic” LLM system settings in hosted environments. In M. Akter, T. Chowdhury, S. Eger, C. Leiter, J. Opitz, & E. Çano (Eds.), Proceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems (pp. 135–148). Association for Computational Linguistics.
Belal, M., She, J., & Wong, S. (2023). Leveraging ChatGPT as text annotation tool for sentiment analysis. arXiv:2306.17177v1.
Bresnan, J., Cueni, A., Nikitina, T., & Baayen, R. H. (2007). Predicting the dative alternation. In G. Bouma, I. Kraemer, & J. Zwarts (Eds.), Cognitive foundations of interpretation (pp. 69–94). KNAW.
Bürkner, P. C. (2017). brms: An R package for Bayesian multilevel models using Stan. Journal of Statistical Software, 80(1), 1–28.
Bybee, J. (2002). Word frequency and context of use in the lexical diffusion of phonetically conditioned sound change. Language Variation and Change, 14(3), 261–290.
Byrt, T., Bishop, J., & Carlin, J. B. (1993). Bias, prevalence and kappa. Journal of Clinical Epidemiology, 46(5), 423–429.
Carpenter, B., Gelman, A., Hoffman, M. D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., & Riddell, A. (2017). Stan: A probabilistic programming language. Journal of Statistical Software, 76(1). 1–32.
Coates, J. (1995). The expression of root and epistemic possibility in English. In J. L. Bybee & S. Fleischman (Eds.), Modality in grammar and discourse (pp. 55–66). John Benjamins.
Cuskley, C., Woods, R., & Flaherty, M. (2024). The limitations of large language models for understanding human language and cognition. Open Mind: Discoveries in Cognitive Science, 8, 1058–1083.
Davies, M. (2002). The Corpus of Historical American English (COHA). English-corpora.org. [URL]
(2016). The News on the Web (NOW) corpus. English-corpora.org. [URL]
Fonteyn, L., Manjavacas, E., & De Regt, J. (2025). Using machine learning to automate data annotation in corpus linguistics: A case study with MacBERTh. International Journal of Corpus Linguistics, 30(3), 296–315.
Fuoli, M., Huang, W., Littlemore, J., Turner, S., & Wilding, E. (2026). Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning. Applied Corpus Linguistics, 6(2), 100204.
Geeraerts, D., Speelman, D., Heylen, K., Montes, M., De Pascale, S., Franco, K., & Lang, M. (2023). Lexical variation and change: A distributional semantic approach. Oxford University Press.
Goldberg, A. E. (1995). Constructions: A construction grammar approach to argument structure. University of Chicago Press.
Gries, S. T., & Berez, A. L. (2017). Linguistic annotation in/for corpus linguistics. In N. Ide & J. Pustejovsky (Eds.), Handbook of linguistic annotation (pp. 379–409). Springer.
Grieve, J., Nini, A., & Guo, D. (2018). Mapping lexical innovation on American social media. Journal of English Linguistics, 46(4), 293–319.
Heylighen, F., & Dewaele, J. -M. (1999). Formality of language: Definition, measurement and behavioral determinants (Internal Report). Center “Leo Apostel”, Free University of Brussels. [URL]
Hoffmann, T., & Trousdale, G. (2011). Variation, change and constructions in English. Cognitive Linguistics, 22(1), 1–23.
Hovy, E., & Lavid, J. (2010). Towards a ‘science’ of corpus annotation: A new methodological challenge for corpus linguistics. International Journal of Translation, 22(1), 13–36.
Jacques, G. (2023). Estimative constructions in cross-linguistic perspective. Linguistic Typology, 27(1), 157–194.
Jaeger, T. F. (2010). Redundancy and reduction: Speakers manage syntactic information density. Cognitive Psychology, 61(1), 23–62.
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv:2001.08361v1.
Kilgarriff, A., Baisa, V., Bušta, J., Jakubíček, M., Kovář, V., Michelfeit, J., Rychlý, P., & Suchomel, V. (2014). The Sketch Engine: Ten years on. Lexicography, 1(1), 7–36.
Leclercq, B., & Morin, C. (2023). No equivalence: A new principle of no synonymy. Constructions, 15(1).
Leivada, E., Dentella, V., & Günther, F. (2024). Evaluating the language abilities of humans vs large language models: Three caveats. Biolinguistics, 18, 1–16.
Levshina, N. (2022). Communicative efficiency: Language structure and use. Cambridge University Press.
Lorenz, D. (2013). Contractions of English semi-modals: The emancipating effect of frequency [Doctoral dissertation, Albert-Ludwigs-Universität Freiburg]. Freiburg University Library.
Luccioni, A. S., Viguier, S., & Ligozat, A. -L. (2023). Estimating the carbon footprint of BLOOM, a 176B parameter language model. The Journal of Machine Learning Research, 24(1), 11990–12004.
Marttinen Larsson, M. (2023). Modelling incipient probabilistic grammar change in real time: The grammaticalisation of possessive pronouns in European Spanish locative adverbial constructions. Corpus Linguistics and Linguistic Theory, 19(2), 177–206.
(2024). Probabilistic reduction and constructionalization: A usage-based diachronic account of the diffusion and conventionalization of the Spanish la de <noun> que construction. Cognitive Linguistics, 35(4), 579–602.
(2025). Pathways of actualization across regional varieties and the real-time dynamics of syntactic change. Language Variation and Change, 37(1), 31–57.
Marttinen Larsson, M., Morin, C., Coussé, E., & Álvarez López, L. (forthcoming). USER-GRAM.
Morin, C., & Marttinen Larsson, M. (2025). Large corpora and large language models: A replicable method for automating grammatical annotation. Linguistics Vanguard, 11(1), 501–510.
Nevalainen, T., & Raumolin-Brunberg, H. (2016). Historical sociolinguistics: Language change in Tudor and Stuart England (2nd ed.). Routledge.
OpenAI. (2025, August 7). Introducing GPT-5. OpenAI.com. [URL]
Ostyakova, L., Mikhailova, A., Molchanova, M., & Konovalov, V. (2025). Redefining annotation practices: Leveraging large language models for discourse annotation. In A. Panchenko, D. Gubanov, M. Khachay, A. Kutuzov, N. Loukachevitch, A. Kuznetsov, I. Nikishina, M. Panov, P. M. Pardalos, A. V. Savchenko, E. Tsymbalov, E. Tutubalina, A. Kasieva, & D. I. Ignatov (Eds.), Analysis of images, social networks and texts: 12th International Conference, AIST 2024, Bishkek, Kyrgyzstan, October 17–19, 2024, revised selected papers (pp. 131–147). Springer.
Ren, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the contrasting narratives on the environmental impact of large language models. Scientific Reports, 14, 26310.
Rickford, J. R., Wasow, T. A., Mendoza-Denton, N., & Espinoza, J. (1995). Syntactic variation and change in progress: Loss of the verbal Coda in topic-restricting as far as constructions. Language, 71(1), 102–131.
Rudnicka, K. (2021). In order that — A data-driven study of symptoms and causes of obsolescence. Linguistics Vanguard, 7(1), 20200092.
Smith, N., Hoffmann, S., & Rayson, P. (2008). Corpus tools and methods, today and tomorrow: Incorporating linguists’ manual annotations. Literary and Linguistic Computing, 23(2), 163–180.
Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In A. Korhonen, D. Traum, & L. Màrquez (Eds.), Proceedings of the 57th annual meeting of the Association for Computational Linguistics (pp. 3645–3650). Association for Computational Linguistics.
The Britannica Dictionary (n.d.) Ask the Editor: “Consider” and “consider as”. Britannica.com. [URL]
Torrent, T. T., Hoffmann, T., Almeida, A. L., & Turner, M. (2024). Copilots for linguists: AI, constructions, and frames. Cambridge University Press.
Yu, D., Li, L., Su, H., & Fuoli, M. (2024). Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology. International Journal of Corpus Linguistics, 29(4), 534–561.