Research Data Leeds Repository

The Corpus of Chinese Academic Written and Spoken English (CAWSE)

Citation

Chen, Yu-Hua, Harrison, Simon, Oakey, David, Stevens, Michael, Yang, Shanru, Ioratim-Uba, Godwin, Zhou, Qianqian and Bruncak, Radovan (2026) The Corpus of Chinese Academic Written and Spoken English (CAWSE). University of Leeds / University of Nottingham NingboChina. [Dataset] https://doi.org/10.5518/1836

Dataset description

The Corpus of Chinese Academic Written and Spoken English (CAWSE) comprises L2 English samples produced by L1 Chinese students enrolled in a preliminary-year programme at an English-medium instruction (EMI) / transnational (TNE) campus in China. Currently available data include coursework (657 scripts; approx. 1,500 tokens per script), interviews (122 sessions, 10 minutes each), and presentations (184 sessions, 10 minutes each), covering the full range of awarded grades. In total, this dataset comprises 657 essays (1,004,523 tokens) and 51 hours of spoken data (352,333 tokens). In addition, a subcorpus of written exam scripts and a multimodal subcorpus of student group discussions may be available upon request. The corpus was compiled between 2016 and 2019, with transcription completed in the following years. It was first released in 2019, and the current release contains only updates to the documentation, with no changes made to the dataset itself. This pre-AI corpus of over one million tokens offers a valuable resource for analysing L2 English, enabling the investigation of lexical, syntactic, and discourse features across proficiency levels, task types, genres, and other contexts. Please refer to Chen et al. (2024) for further details about the corpus.

Additional information: This dataset is held in the Restricted Access Data Repository, RADAR. To request access, please click the link under ‘Related resources’. You will be prompted to complete a request form including your intended use of the dataset. You will then receive an initial response to your request within 10 working days.
Keywords: corpus; EMI; academic English; L2 English; L1 Chinese
Subjects: Q000 - Linguistics, classics & related subjects > Q100 - Linguistics > Q110 - Applied linguistics
Related resources:
LocationType
https://radar.researchdata.leeds.ac.uk/81/Dataset
Date deposited: 01 Jul 2026 09:24
URI: https://archive.researchdata.leeds.ac.uk/id/eprint/1568

Files

Documentation

Research Data Leeds Repository is powered by EPrints
Copyright © University of Leeds