Article published In: International Journal of Learner Corpus Research: Online-First Articles
The KSAUHS learner corpus
A longitudinal resource on Arabic L1 EFL academic writing
Eman Al Nafjan | King Saud bin Abdulaziz University for Health Sciences | King Abdullah International Medical Research Centre | Ministry of the National Guard — Health Affairs
Alaa Alfelaij | King Saud bin Abdulaziz University for Health Sciences | King Abdullah International Medical Research Centre | Ministry of the National Guard — Health Affairs
Norah Alfawaz | King Saud bin Abdulaziz University for Health Sciences | King Abdullah International Medical Research Centre | Ministry of the National Guard — Health Affairs
Sarah Mohammed | King Saud bin Abdulaziz University for Health Sciences | King Abdullah International Medical Research Centre | Ministry of the National Guard — Health Affairs
Nouf Ali Binsuwadan | King Saud bin Abdulaziz University for Health Sciences | King Abdullah International Medical Research Centre | Ministry of the National Guard — Health Affairs
Published online: 17 July 2026
https://doi.org/10.1075/ijlcr.25028.aln
https://doi.org/10.1075/ijlcr.25028.aln
Abstract
The KSAU-HS Learner Corpus is a longitudinal corpus of EFL tertiary writing that complies with the FAIR
principles. Collection began in 2022 and captures writing development during a period of emerging language technologies (2022–24).
The corpus contains over 856,907 tokens across 2,387 texts produced by 157 preparatory year university students, within a
CEFR-aligned program with an instructional range of approximately A2-B2. Texts span four trimesters and include rhetorical modes
such as cause-and-effect, argumentation, and summarisation. Metadata includes years of English schooling, other languages spoken,
and preferred reference tools. The corpus enables research into writing development to inform EAP pedagogy and assessment. Avenues
for investigation include lexico-grammatical development, cross-linguistic influence, individual differences, and the impact of
task conditions and language technologies. This resource promises data-driven insights into the textual features and factors that
characterise EFL writing proficiency, and plans are in place to expand its size, representation, and accessibility.
Article outline
- 1.Introduction
- 2.Key features of the learner corpus
- 2.1Longitudinal design
- 2.2Extensive metadata
- 2.3Discipline-linked proficiency differentiation
- 2.4Register diversity and task variation
- 2.5Temporally situated around generative AI emergence
- 2.6Continued access to participants for follow-up
- 2.7Alignment with FAIR Principles
- 3.Corpus design and composition
- 4.Writing task progression across trimesters
- 5.Workflow and annotation procedures
- 6.Corpus composition across subcorpora
- 7.Ethical considerations
- 8.Future directions and conclusion
- Open data badge and data availability statement
- AI use disclosure statement
- Acknowledgements
References
References (32)
Alfuraih, R. F. (2020). The
undergraduate learner translator corpus: A new resource for translation studies and computational
linguistics. Language Resources and
Evaluation, 54(3), 801–830.
Algouzi, S. (2021). Functions
of the discourse marker so in the LINDSEI-AR corpus. Cogent Arts &
Humanities, 8(1), 1872166.
Al-Harthi, M., Alsaif, A., Al-Nafjan, E., Alshihri, F., & Saleh, M. (2024). Saudi
Learner Translation Corpus: The design and compilation of an English — Arabic learner translation
corpus. PLoS
ONE, 19(10), e0303729.
Al Nafjan, E., & Jawhar, S. (2025). Learner
corpus research and data-driven learning in Saudi Arabia. In A. H. Al-Hoorie, C. Mitchell, & T. Elyas (Eds.), Language
education in Saudi Arabia: Integrating technology in the
classroom. Springer.
Althewini, A., Alkushi, A., Alhawsawi, S. Y., Al. Roomy, M., Jawhar, S. S., Aldafas, A., Almusaad, H. R., & Alnafisah, M. (2025). Examining
the influence of admission criteria and English grades on performance in science and humanities courses at a Saudi medical
university. Eurasian Journal of Educational
Research, 116. [URL]
Biber, D., Conrad, S., & Reppen, R. (1998). Corpus
linguistics: Investigating language structure and use. Cambridge University Press.
Biber, D., Reppen, R., Staples, S., & Egbert, J. (2020). Exploring
the longitudinal development of grammatical complexity in the disciplinary writing of L2 English university
students. International Journal of Learner Corpus
Research, 6(1), 38–71.
Boulton, A., & Vyatkina, N. (2021). Thirty
years of data-driven learning: Taking stock and charting new directions over time. Language
Learning &
Technology, 25(3), 66–89.
Callies, M. (2023). Current
perspectives on learner corpus research. AAA: Arbeiten aus Anglistik und
Amerikanistik, 48(1), 37–52. [URL]
Carlsen, C. (2012). Proficiency
level — A fuzzy variable in computer learner corpora, Applied
Linguistics, 33(2), 161–183.
Cheung, L., & Crosthwaite, P. (2025). CorpusChat:
Integrating corpus linguistics and generative AI for academic writing development. Computer
Assisted Language Learning, 1–27.
Deshors, S. C., & Gries, S. T. (2020). Comparing
learner corpora. In N. Tracy-Ventura, & M. Paquot (Eds.), The
Routledge handbook of SLA and
corpora (pp. 107–120). Routledge.
Díaz-Negrillo, A., & Thompson, P. (2013). Automatic
annotation and analysis of learner corpus
data. Amsterdam: John Benjamins.
Forti, L. (2024). Proficiency-rated
learner corpora: A promising resource for data-driven learning. International Journal of
Learner Corpus Research
(IJLCR), 10(1), 216–240. [URL].
Gilquin, G. (2008). Hesitation
markers among EFL learners: Pragmatic deficiency or
difference? In J. Romero-Trillo (Ed.), Pragmatics
and corpus linguistics: A mutualistic
entente (Vol. 21, pp. 119–136). Université catholique de Louvain. [URL].
Granger, S. (2002). A
bird’s-eye view of learner corpus research. In S. Granger, J. Hung, & S. Petch-Tyson (Eds.), Computer
learner corpora, Second language acquisition and foreign language
teaching (pp. 3–33). Amsterdam: John Benjamins.
(2024). From
early to future learner corpus research. International Journal of Learner Corpus
Research, 10(2), 247–279.
Granger, S., Dupont, M., Meunier, F., Naets, H., & Paquot, M. (2020). The
International Corpus of Learner English (Version 3). Presses universitaires de Louvain.
Granger, S., & Lefer, M. A. (2020). The
Multilingual Student Translation corpus: a resource for translation teaching and
research. Language Resources and
Evaluation, 541, 1183–1199.
Habash, N., & Palfreyman, D. (2022). ZAEBUC:
An annotated Arabic-English bilingual writer corpus. In N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, & S. Piperidis (Eds.), Proceedings
of the thirteenth language resources and evaluation
conference (pp. 79–88). European Language Resources Association.
Ishikawa, S. (2023). The
ICNALE guide: An introduction to a learner corpus study on Asian learners’ L2 English (1st
ed.). Routledge.
Larsson, T., Paquot, M., & Biber, D. (2021). On
the importance of register in learner writing. In E. Seoane & D. Biber (Eds.), Studies
in corpus
linguistics (Vol. 1031, pp. 235–258). John Benjamins.
McEnery, T., & Hardie, A. (2012). Corpus
linguistics: Method, theory and practice. Cambridge University Press.
The Regents of the University of
Michigan. (2009). Michigan Corpus of Upper-level Student
Papers. (2009). The Regents of the University of Michigan. [URL]
Meunier, F. (2016). Introduction
to the LONGDALE project. In E. Castello, K. Ackerley, & F. Coccetta (Eds.), Studies
in learner corpus linguistics: Research and applications for foreign language teaching and
assessment (pp. 123–126). Peter Lang.
Myles, F. (2021). Commentary:
An SLA perspective on learner corpus research. In B. Le Bruyn, & M. Paquot (Eds.), Learner
corpus research meets second language
acquisition (pp. 258–273). Cambridge University Press.
Naismith, B., Han, N. R., & Juffs, A. (2022). The
University of Pittsburgh English Language Institute Corpus (PELIC). International Journal of
Learner Corpus
Research, 8(1), 121–138.
Nesi, H., & Gardner, S. (2012). Genres
across the disciplines: Student writing in higher education. Cambridge University Press.
Paquot, M. (2024). Learner
corpus research: A critical appraisal and roadmap for contributing (more) to SLA research
agendas. Corpus Linguistics and Linguistic
Theory, 20(3), 567–590.
Paquot, M., König, A., Stemle, E. W., & Frey, J. C. (2024). The
core metadata schema for learner corpora (LC-meta): Collaborative efforts to advance data discoverability, metadata quality
and study comparability in L2 research. International Journal of Learner Corpus
Research, 10(2), 280–300.
Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J. W., Bonino da Silva Santos, L., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo., C. T., Finkers, R., … & Mons, B. (2016). The
FAIR guiding principles for scientific data management and stewardship. Scientific
Data, 3(1), 160018.