How verbal and non-verbal behaviours shape collaborative learning processes: Towards multimodal evidence for socially shared regulation
Thesis event information
Date and time of the thesis defence
Place of the thesis defence
Pohjanmaansali (L1), Linnanmaa campus
Topic of the dissertation
How verbal and non-verbal behaviours shape collaborative learning processes: Towards multimodal evidence for socially shared regulation
Doctoral candidate
Master of Arts (Education) Ridwan Whitehead
Faculty and unit
University of Oulu Graduate School, Faculty of Education and Psychology, Learning and Educational Technology Research Lab
Subject of study
Educational sciences
Opponent
Professor Guido Makransky, University of Copenhagen
Custos
Professor Sanna Järvelä, University of Oulu
Gazes and physical distance shape how groups notice and work through challenges in collaborative learning
Collaboration is increasingly important in education and working life, and group work is widely used to support learning. However, productive collaboration does not happen simply by placing learners in groups. Working together is cognitively, emotionally, and motivationally demanding, and groups regularly face challenges such as misunderstandings, loss of motivation, or tension between members. Successful groups notice these moments and adjust their joint work, a process known as socially shared regulation of learning. So far, research has studied this mainly through what learners say, even though much of human interaction happens without words.
This dissertation examines how socially shared regulation begins and is coordinated through the interplay of verbal and non-verbal behaviour during collaborative learning. In particular, it explores how gaze, posture, and the physical distance between group members are intertwined with the ways groups recognise challenges and respond to them together.
The dissertation consists of three studies whose data were collected in settings ranging from a controlled task to an authentic classroom. In the first study, upper secondary school students worked in small groups on a task with planned cognitive and emotional challenges, and their gaze and talk were analysed from video and audio recordings. In the second study, pre-service teachers completed a collaborative physics task, and the video was used to test whether AI models could recognise their postures. In the third study, 12- to 14-year-old pupils worked in groups during AI-supported science lessons, and a computer vision system automatically measured from ordinary classroom video where pupils looked and how far apart they were.
The results show that gaze and talk unfold together in a meaningful order. For example, planning discussions were followed by students turning their gaze to the task, and looking at peers was followed by talk about feelings and the group atmosphere. How often students looked at each other did not change when challenges arose, suggesting that gaze keeps the group attentive and ready to act rather than signalling on its own when regulation starts. In the classroom, eye contact between group members often came before or alongside the moment the group began checking how their work was going, making emerging problems visible, sometimes before anyone voiced them. Physical distance also mattered. When group members sat further apart, they had more difficulties organising their work and seemed to compensate by looking at each other more. When they sat closer, shared focus on the same materials helped them move from noticing a problem to choosing and carrying out a strategy.
Methodologically, the dissertation shows how body language can be extracted from video in reliable and theory-informed ways. AI models that interpret images produced highly consistent posture labels, but their agreement with human coders varied across posture types, showing that automated tools must be tested like any other measuring instrument. The computer vision approach made it possible to measure gaze and distance in a real classroom without disrupting natural interaction.
Overall, the dissertation deepens understanding of collaborative learning by showing that groups regulate their learning not only through words but also through where group members look and how they position themselves in relation to one another. For educational practice, the findings highlight the value of shared focal points such as common screens or visible task materials, attention to classroom layout and seating, and teachers noticing changes in group members' attention as possible early signs of difficulties. The results also offer guidance for developing responsible, privacy-conscious tools that support teachers and learners during collaborative work.
This dissertation examines how socially shared regulation begins and is coordinated through the interplay of verbal and non-verbal behaviour during collaborative learning. In particular, it explores how gaze, posture, and the physical distance between group members are intertwined with the ways groups recognise challenges and respond to them together.
The dissertation consists of three studies whose data were collected in settings ranging from a controlled task to an authentic classroom. In the first study, upper secondary school students worked in small groups on a task with planned cognitive and emotional challenges, and their gaze and talk were analysed from video and audio recordings. In the second study, pre-service teachers completed a collaborative physics task, and the video was used to test whether AI models could recognise their postures. In the third study, 12- to 14-year-old pupils worked in groups during AI-supported science lessons, and a computer vision system automatically measured from ordinary classroom video where pupils looked and how far apart they were.
The results show that gaze and talk unfold together in a meaningful order. For example, planning discussions were followed by students turning their gaze to the task, and looking at peers was followed by talk about feelings and the group atmosphere. How often students looked at each other did not change when challenges arose, suggesting that gaze keeps the group attentive and ready to act rather than signalling on its own when regulation starts. In the classroom, eye contact between group members often came before or alongside the moment the group began checking how their work was going, making emerging problems visible, sometimes before anyone voiced them. Physical distance also mattered. When group members sat further apart, they had more difficulties organising their work and seemed to compensate by looking at each other more. When they sat closer, shared focus on the same materials helped them move from noticing a problem to choosing and carrying out a strategy.
Methodologically, the dissertation shows how body language can be extracted from video in reliable and theory-informed ways. AI models that interpret images produced highly consistent posture labels, but their agreement with human coders varied across posture types, showing that automated tools must be tested like any other measuring instrument. The computer vision approach made it possible to measure gaze and distance in a real classroom without disrupting natural interaction.
Overall, the dissertation deepens understanding of collaborative learning by showing that groups regulate their learning not only through words but also through where group members look and how they position themselves in relation to one another. For educational practice, the findings highlight the value of shared focal points such as common screens or visible task materials, attention to classroom layout and seating, and teachers noticing changes in group members' attention as possible early signs of difficulties. The results also offer guidance for developing responsible, privacy-conscious tools that support teachers and learners during collaborative work.
Created 11.9.2026 | Updated 14.9.2026