Corpus Linguistics
A.Y. 2026/2027
Learning objectives
The course aims to provide students with a solid introduction to the theoretical, methodological, and applied principles of corpus linguistics, adopting a cross-linguistic approach suited to different languages of study. Through three progressive modules, the course seeks to:
- Introduce the fundamental concepts of corpus linguistics;
- Provide skills for collecting, querying, and analysing corpus-based data;
- Develop critical thinking skills in linguistic research contexts, supported by practical examples applicable across different languages.
- Introduce the fundamental concepts of corpus linguistics;
- Provide skills for collecting, querying, and analysing corpus-based data;
- Develop critical thinking skills in linguistic research contexts, supported by practical examples applicable across different languages.
Expected learning outcomes
By the end of the course, students will have acquired a solid understanding of the theoretical and methodological foundations of corpus linguistics, including its core principles and the limitations associated with the use of data extracted from corpora. They will be familiar with the main operational procedures for extracting and organizing linguistic data and will be able to apply quantitative analysis methods, having developed the skills needed to interpret fundamental statistical measures in linguistics.
Throughout the course activities, students will learn how to use specific software for querying and analysing corpora, how to design simple empirical investigations, and how to apply corpus-based tools to linguistic data from different languages. Particular attention will be given to developing the critical thinking skills necessary to evaluate corpus data, reflect on methodological choices, and understand the potential and limitations of these tools in language analysis.
Thanks to the knowledge acquired and the development of critical thinking fostered during the course, students will be able to independently pursue further study of corpus-based methodologies applied to their chosen language of specialisation.
Throughout the course activities, students will learn how to use specific software for querying and analysing corpora, how to design simple empirical investigations, and how to apply corpus-based tools to linguistic data from different languages. Particular attention will be given to developing the critical thinking skills necessary to evaluate corpus data, reflect on methodological choices, and understand the potential and limitations of these tools in language analysis.
Thanks to the knowledge acquired and the development of critical thinking fostered during the course, students will be able to independently pursue further study of corpus-based methodologies applied to their chosen language of specialisation.
Lesson period: First semester
Assessment methods: Esame
Assessment result: voto verbalizzato in trentesimi
Single course
This course can be attended as a single course.
Course syllabus and organization
Single session
Responsible
Lesson period
First semester
Course syllabus
The course is structured into three main modules: A, B, and C.
Module A - Introduction to Corpus Linguistics
Module A provides a general introduction to corpus linguistics, presenting its main theoretical, methodological and operational foundations. The module will discuss the criteria that define a linguistic corpus and will introduce key issues involved in corpus design, description and use.
Attention will be given to the relationship between corpora, research questions and linguistic evidence, with reference to different types of corpora and to some of the methodological choices involved in corpus-based research. The module will also provide an overview of the development of corpus linguistics as a discipline and of its role in contemporary language studies.
An operational introduction to Sketch Engine will also be included. Interactive activities in Module A will guide students in observing linguistic data, comparing intuition and evidence, reflecting on corpus definitions and beginning to use corpus tools in a guided way.
Module B - Quantitative Methods for Corpus Data Analysis
Module B introduces students to the main conceptual, quantitative and operational tools for analysing data derived from linguistic corpora. The module presents basic notions such as token, type and lemma, lexical variety, type-token ratio, standardised and moving average type-token ratio, absolute and relative frequency, and normalisation.
The module also introduces the distribution of frequencies, Zipf's Law, the notions of collocations, node and collocate, span, bigrams, observed and expected frequency, and the main association measures used to describe co-occurrences between linguistic items, including T-score, Mutual Information and LogDice. The concept of keyness will also be addressed, with particular attention to the relationship between focus corpus and reference corpus, keywords and terms.
The aim of the module is to equip students with the tools necessary to understand and critically interpret quantitative results produced by corpus analysis software.
Interactive activities in Module B include guided exercises and inductive tasks. Students will work on token/type counts, lexical variety, relative frequency and normalisation, wordlists and rank/frequency graphs, collocations, association measures and keyness.
Module C - Practical Applications
Module C is dedicated to the application of corpus linguistics principles and methods. The module focuses on selected guided activities and mini-project work designed to consolidate the analytical skills developed in the previous modules.
The module introduces students to the formulation of research questions that are compatible with a corpus-based approach, that is, questions that can be addressed through the analysis of linguistic data observable in a corpus. Students will be guided to reflect on the relationship between research question, corpus selection, data type, analytical method and limits of interpretation.
The goal of the module is to foster the ability to apply corpus linguistics tools critically, while keeping the analysis anchored to observable linguistic evidence.
Module A - Introduction to Corpus Linguistics
Module A provides a general introduction to corpus linguistics, presenting its main theoretical, methodological and operational foundations. The module will discuss the criteria that define a linguistic corpus and will introduce key issues involved in corpus design, description and use.
Attention will be given to the relationship between corpora, research questions and linguistic evidence, with reference to different types of corpora and to some of the methodological choices involved in corpus-based research. The module will also provide an overview of the development of corpus linguistics as a discipline and of its role in contemporary language studies.
An operational introduction to Sketch Engine will also be included. Interactive activities in Module A will guide students in observing linguistic data, comparing intuition and evidence, reflecting on corpus definitions and beginning to use corpus tools in a guided way.
Module B - Quantitative Methods for Corpus Data Analysis
Module B introduces students to the main conceptual, quantitative and operational tools for analysing data derived from linguistic corpora. The module presents basic notions such as token, type and lemma, lexical variety, type-token ratio, standardised and moving average type-token ratio, absolute and relative frequency, and normalisation.
The module also introduces the distribution of frequencies, Zipf's Law, the notions of collocations, node and collocate, span, bigrams, observed and expected frequency, and the main association measures used to describe co-occurrences between linguistic items, including T-score, Mutual Information and LogDice. The concept of keyness will also be addressed, with particular attention to the relationship between focus corpus and reference corpus, keywords and terms.
The aim of the module is to equip students with the tools necessary to understand and critically interpret quantitative results produced by corpus analysis software.
Interactive activities in Module B include guided exercises and inductive tasks. Students will work on token/type counts, lexical variety, relative frequency and normalisation, wordlists and rank/frequency graphs, collocations, association measures and keyness.
Module C - Practical Applications
Module C is dedicated to the application of corpus linguistics principles and methods. The module focuses on selected guided activities and mini-project work designed to consolidate the analytical skills developed in the previous modules.
The module introduces students to the formulation of research questions that are compatible with a corpus-based approach, that is, questions that can be addressed through the analysis of linguistic data observable in a corpus. Students will be guided to reflect on the relationship between research question, corpus selection, data type, analytical method and limits of interpretation.
The goal of the module is to foster the ability to apply corpus linguistics tools critically, while keeping the analysis anchored to observable linguistic evidence.
Prerequisites for admission
No specific prior knowledge is required, apart from a basic understanding of general linguistics.
Teaching methods
The course is organised as an online course and combines synchronous and asynchronous didactic delivery with synchronous and asynchronous interactive teaching activities.
Asynchronous didactic delivery will consist of short, self-contained video lectures. Synchronous didactic delivery will be used to introduce and discuss selected topics, especially those dealing with quantitative methods of analysis.
Interactive teaching activities will include guided exercises, practical work on corpus data, use of Sketch Engine, short written outputs, web labs and synchronous feedback sessions. These activities are designed to support students in applying the concepts introduced in the course and in developing the ability to interpret corpus data critically.
Asynchronous didactic delivery will consist of short, self-contained video lectures. Synchronous didactic delivery will be used to introduce and discuss selected topics, especially those dealing with quantitative methods of analysis.
Interactive teaching activities will include guided exercises, practical work on corpus data, use of Sketch Engine, short written outputs, web labs and synchronous feedback sessions. These activities are designed to support students in applying the concepts introduced in the course and in developing the ability to interpret corpus data critically.
Teaching Resources
All teaching materials required for the course will be made available on myAriel. These materials will include slides, handouts, guided activities, practical exercises, selected corpus outputs and instructions for the use of Sketch Engine.
The references listed below are intended as optional further reading for students who wish to explore specific topics in greater depth:
Freddi, Maria (2019). Linguistica dei corpora. Roma: Carocci.
Barbera, M. (2013). Linguistica dei corpora e linguistica dei corpora italiana. Una introduzione. Milano: Quasar.
McEnery, T. and Wilson, A. (2001). Corpus Linguistics: An Introduction. 2nd Edition. Edinburgh: Edinburgh University Press.
Brezina, V. (2018). Statistics in Corpus Linguistics: A Practical Guide. Cambridge University Press.
The syllabus is the same for attending and non-attending students.
The references listed below are intended as optional further reading for students who wish to explore specific topics in greater depth:
Freddi, Maria (2019). Linguistica dei corpora. Roma: Carocci.
Barbera, M. (2013). Linguistica dei corpora e linguistica dei corpora italiana. Una introduzione. Milano: Quasar.
McEnery, T. and Wilson, A. (2001). Corpus Linguistics: An Introduction. 2nd Edition. Edinburgh: Edinburgh University Press.
Brezina, V. (2018). Statistics in Corpus Linguistics: A Practical Guide. Cambridge University Press.
The syllabus is the same for attending and non-attending students.
Assessment methods and Criteria
The exam consists of a written test including both theoretical questions and practical exercises, aimed at assessing students' understanding of the course content and their ability to apply the methodologies covered in the three course modules.
Modules or teaching units
Part A and B
GLOT-01/A - Historical and General Linguistics - University credits: 6
Lessons: 40 hours
Professor:
Berti Barbara
Part C
GLOT-01/A - Historical and General Linguistics - University credits: 3
Lessons: 20 hours
Professor:
Berti Barbara
Professor(s)