Research Data Leeds Repository
The Corpus of Chinese Academic Written and Spoken English (CAWSE)
Citation
Chen, Yu-Hua, Harrison, Simon, Oakey, David, Stevens, Michael, Yang, Shanru, Ioratim-Uba, Godwin, Zhou, Qianqian and Bruncak, Radovan (2026) The Corpus of Chinese Academic Written and Spoken English (CAWSE). University of Leeds / University of Nottingham NingboChina. [Dataset] https://doi.org/10.5518/1836
Dataset description
The Corpus of Chinese Academic Written and Spoken English (CAWSE) comprises L2 English samples produced by L1 Chinese students enrolled in a preliminary-year programme at an English-medium instruction (EMI) / transnational (TNE) campus in China. Currently available data include coursework (657 scripts; approx. 1,500 tokens per script), interviews (122 sessions, 10 minutes each), and presentations (184 sessions, 10 minutes each), covering the full range of awarded grades. In total, this dataset comprises 657 essays (1,004,523 tokens) and 51 hours of spoken data (352,333 tokens). In addition, a subcorpus of written exam scripts and a multimodal subcorpus of student group discussions may be available upon request. The corpus was compiled between 2016 and 2019, with transcription completed in the following years. It was first released in 2019, and the current release contains only updates to the documentation, with no changes made to the dataset itself. This pre-AI corpus of over one million tokens offers a valuable resource for analysing L2 English, enabling the investigation of lexical, syntactic, and discourse features across proficiency levels, task types, genres, and other contexts. Please refer to Chen et al. (2024) for further details about the corpus.
| Additional information: | This dataset is held in the Restricted Access Data Repository, RADAR. To request access, please click the link under ‘Related resources’. You will be prompted to complete a request form including your intended use of the dataset. You will then receive an initial response to your request within 10 working days. | ||||
|---|---|---|---|---|---|
| Keywords: | corpus; EMI; academic English; L2 English; L1 Chinese | ||||
| Subjects: | Q000 - Linguistics, classics & related subjects > Q100 - Linguistics > Q110 - Applied linguistics | ||||
| Related resources: |
|
||||
| Date deposited: | 01 Jul 2026 09:24 | ||||
| URI: | https://archive.researchdata.leeds.ac.uk/id/eprint/1568 | ||||



CAWSE Readme (2026) [508kB]
CAWSE Readme (2026) [508kB]