Corpus Linguistics

A.Y. 2026/2027
9
Max ECTS
60
Overall hours
SSD
GLOT-01/A
Language
Italian
Learning objectives
The course aims to provide students with a solid introduction to the theoretical, methodological, and applied principles of corpus linguistics, adopting a cross-linguistic approach suited to different languages of study. Through three progressive modules, the course seeks to:
- Introduce the fundamental concepts of corpus linguistics;
- Provide skills for collecting, querying, and analysing corpus-based data;
- Develop critical thinking skills in linguistic research contexts, supported by practical examples applicable across different languages.
Expected learning outcomes
By the end of the course, students will have acquired a solid understanding of the theoretical and methodological foundations of corpus linguistics, including its core principles and the limitations associated with the use of data extracted from corpora. They will be familiar with the main operational procedures for extracting and organizing linguistic data and will be able to apply quantitative analysis methods, having developed the skills needed to interpret fundamental statistical measures in linguistics.
Throughout the course activities, students will learn how to use specific software for querying and analysing corpora, how to design simple empirical investigations, and how to apply corpus-based tools to linguistic data from different languages. Particular attention will be given to developing the critical thinking skills necessary to evaluate corpus data, reflect on methodological choices, and understand the potential and limitations of these tools in language analysis.
Thanks to the knowledge acquired and the development of critical thinking fostered during the course, students will be able to independently pursue further study of corpus-based methodologies applied to their chosen language of specialisation.
Single course

This course can be attended as a single course.

Course syllabus and organization

Single session

Responsible
Lesson period
First semester
Course syllabus
The course is structured into three parts: A, B and C.

Part A - Introduction to Corpus Linguistics
Part A provides a general introduction to Corpus Linguistics, presenting its main theoretical, methodological and practical foundations. The module will discuss the criteria that define a linguistic corpus and introduce key issues relating to corpus design, description and use.
Particular attention will be paid to the relationship between corpora, research questions and linguistic evidence, with reference to different types of corpora and some of the methodological choices involved in corpus-based research. The module will also provide an overview of the development of Corpus Linguistics as a discipline and its role in contemporary linguistic research.
The module will include a practical introduction to Sketch Engine. The interactive activities in Part A will guide students in observing linguistic data, comparing intuition with empirical evidence, reflecting on definitions of a corpus, and taking their first guided steps in the use of corpus-based analysis tools.

Part B - Quantitative Methods for Linguistic Data Analysis
Part B introduces students to the main conceptual, quantitative and practical tools used to analyse data extracted from linguistic corpora. The module presents basic concepts such as token, type and lemma; lexical diversity; type-token ratio, standardised type-token ratio and moving-average type-token ratio; absolute and relative frequency; and normalisation.
The module also introduces frequency distributions, Zipf's Law, and the concepts of collocation, node and collocate, span, bigrams, observed frequency and expected frequency. It also covers the main association measures used to describe co-occurrences between linguistic items, including T-score, Mutual Information and LogDice. The concept of keyness will also be addressed, with particular attention to the relationship between a focus corpus and a reference corpus, as well as to keywords and terms.
The aim of the module is to provide students with the tools needed to understand and critically interpret the quantitative results produced by corpus analysis software.
The interactive activities in Part B include guided exercises and inductive tasks. Students will work on token and type counts, lexical diversity, relative frequency and normalisation, wordlists and rank-frequency plots, collocations, association measures and keyness.

Part C - Practical Applications
Part C is devoted to the application of the principles and methods of Corpus Linguistics. The module focuses on selected guided activities and mini-projects designed to consolidate the analytical skills developed in the previous parts.
Part C introduces students to the formulation of research questions suitable for a corpus-based approach. Students will be guided in reflecting on the relationship between the research question, corpus selection, data type, analytical method and the limitations of interpretation.
The aim of the module is to develop students' ability to apply the tools of Corpus Linguistics critically, while ensuring that their analyses remain grounded in observable linguistic evidence.
Prerequisites for admission
No specific prior knowledge is required, apart from a basic understanding of general linguistics.
Teaching methods
The course is organised as an online course and combines synchronous and asynchronous didactic delivery with interactive teaching activities.
Asynchronous didactic delivery will consist of short, self-contained video lectures. Synchronous didactic delivery will be used to introduce and discuss selected topics, especially those dealing with quantitative methods of analysis.
Interactive teaching activities will include guided exercises, practical work, use of Sketch Engine, etc. These activities are designed to support students in applying the concepts introduced in the course and in developing the ability to interpret corpus data critically.
Teaching Resources
All teaching materials required for the course will be made available on myAriel. These materials will include slides, guided activities, practical exercises, selected corpus outputs and instructions for the use of Sketch Engine.

The references listed below are intended as optional further reading for students who wish to explore specific topics in greater depth:

Freddi, Maria (2019). Linguistica dei corpora. Roma: Carocci.
Barbera, M. (2013). Linguistica dei corpora e linguistica dei corpora italiana. Una introduzione. Milano: Quasar.
McEnery, T. and Wilson, A. (2001). Corpus Linguistics: An Introduction. 2nd Edition. Edinburgh: Edinburgh University Press.
Brezina, V. (2018). Statistics in Corpus Linguistics: A Practical Guide. Cambridge University Press.

The syllabus is the same for attending and non-attending students.
Assessment methods and Criteria
The exam consists of a written test including both theoretical questions and practical exercises, aimed at assessing students' understanding of the course content and their ability to apply the methodologies covered in the three course modules.
Modules or teaching units
Part A and B
GLOT-01/A - Historical and General Linguistics - University credits: 6
Lessons: 40 hours
Professor: Berti Barbara

Part C
GLOT-01/A - Historical and General Linguistics - University credits: 3
Lessons: 20 hours
Professor: Berti Barbara

Educational website(s)
Professor(s)