Back to portfolioVolver al portfolio Gerben Suerink.
UX research, Behavioural psychology, Field study, Human-AI interaction, TU/e thesisInvestigación UX, Psicología conductual, Estudio de campo, Interacción humano-IA, Tesis TU/e

When interfaces watch you workCuando las interfaces te miran trabajar

A graduation research study on attention-monitoring interfaces. When the tool can see whether the user is paying attention, how does live feedback change their performance, their self-perception, and their relationship with the tool itself? Two framing conditions, a between-subjects design, eye-tracker telemetry. The finding that travels furthest: how you introduce an AI feedback system shapes reactance more than its accuracy does, and how often it interrupts shapes reactance more than how you introduce it.

Un estudio de investigación de grado sobre interfaces que monitorizan la atención. Cuando la herramienta puede ver si el usuario está prestando atención, ¿cómo cambia el feedback en vivo su rendimiento, su autopercepción y su relación con la propia herramienta? Dos condiciones de encuadre, un diseño entre sujetos, telemetría de eye-tracker. El hallazgo que más viaja: cómo introduces un sistema de feedback con IA da forma a la reactancia más que su precisión real, y la frecuencia con la que interrumpe da forma a la reactancia más que cómo lo introduces.

In partnership withEn colaboración con
Eindhoven University of Technology (TU/e)

Department of Industrial Engineering and Innovation Sciences · Bachelor End Project · June 2026 · Sole-authored thesis, supervised by Linghan Zhang. Lab infrastructure shared with two parallel BEPs.

Departamento de Ingeniería Industrial y Ciencias de la Innovación · Proyecto de Grado · junio de 2026 · Tesis de autoría única, supervisada por Linghan Zhang. Infraestructura de laboratorio compartida con dos proyectos paralelos.

My RoleMi rol

Sole author. Designed the framing study, ran the field sessions, analysed the results, wrote the thesis. Bachelor End Project at TU/e under the Psychology & Technology programme, supervised by Linghan Zhang. Five-month project from literature scan to publication. Lab infrastructure was shared with two parallel BEPs (each studying a different manipulation within the same 2×2 design); this report and its framing manipulation are mine.

Autor único. Diseñé el estudio de encuadre, llevé las sesiones de campo, analicé los resultados y escribí la tesis. Proyecto de Grado en la TU/e dentro del programa de Psicología y Tecnología, supervisado por Linghan Zhang. Proyecto de cinco meses, desde el barrido de literatura hasta la publicación. La infraestructura de laboratorio se compartió con dos proyectos paralelos (cada uno estudiando una manipulación diferente dentro del mismo diseño 2×2); este informe y su manipulación del encuadre son míos.

Problem Statement

Attention-monitoring interfaces (eye-tracker focus detectors, productivity sensors) are entering everyday products. They promise self-improvement, but previous experiments have repeatedly failed to translate technical accuracy into actual learning gain.

The gap is psychological, not technical. Does the way you introduce the system to the user change how the same system performs?

Research Goals

  • Measure how framing depth (brief vs in-depth) shifts perceived accuracy.
  • Measure how perceived accuracy translates to reactance toward the system.
  • Measure whether reactance reduction translates to actual learning gain.
  • Surface design implications for products shipping AI feedback systems.

Constraints

  • TU/e ethics approval boundaries on biometric data.
  • 5-month timeline (literature, design, field, analysis, write-up).
  • Shared lab budget with two parallel BEPs (joint 2×2 design).
  • Recruitment cap from TU/e participant pool.
  • No commercial budget. Open-source tooling only.
Study at a glance
5 moend-to-end project
2framing conditions tested
4measured outcomes per session
N=15analysis sample (after exclusions)
d = +2.00
framing effect on user reactance (p = .008)

Part I: Literature

A scan across HCI literature on attention monitoring, trust in algorithmic interfaces, psychological reactance, and cognitive load. Three findings shaped the design.

Anchor 1Meaningful explanations enhance perceived system accuracy independently of actual performance (Nourani et al., 2019).
Anchor 2Psychological reactance theory: forced engagement creates resistance, consuming cognitive resources otherwise available for learning (Brehm, 1966; Sweller et al., 1998).
Anchor 3Prior attention-monitoring studies restored attention without producing learning gains (Robal et al., 2018; Ye, 2024). The pedagogical effectiveness gap is the gap this study targets.

Part II: Study design

A between-subjects design conducted inside a shared 2×2 experimental infrastructure with two parallel BEPs. The present thesis isolates the framing manipulation by holding pop-up type constant. Participants were randomly assigned to a brief framing condition (~70 words, minimal introductory text about the monitoring system) or an in-depth framing condition (~200 words, additionally enumerating the four eye-behaviour metrics contributing to the attention score). Tone, supportive language, and explicit accuracy claims were equalised across conditions, leaving depth of mechanistic disclosure as the manipulated feature.

Participants then watched a 13.5-minute lecture video while a Tobii eye-tracker measured attention. When the attention score fell below threshold, a pop-up paused the video. After the lecture, a post-lecture quiz (11 multiple-choice items) measured learning performance, followed by validated scales for perceived accuracy, perceived intrusiveness, and psychological reactance (PRS-HCI).

BETWEEN-SUBJECTS FRAMING MANIPULATION BRIEF FRAMING ~70 words "Eye-tracker monitors attention." IN-DEPTH FRAMING ~200 words 4 metrics: fixation, gaze, blink, inactivity IDENTICAL LECTURE + MONITORING 13.5-min lecture · Tobii eye-tracker · attention pop-ups FOUR MEASURED OUTCOMES • Perceived system accuracy (5-item, α=.945) • Perceived intrusiveness (7-item, α=.891) • Psychological reactance, PRS-HCI (8-item, α=.909) • Learning performance (11-item post-lecture quiz)
Between-subjects design. The only thing varying between groups was the introductory text.

Decision 1: How to isolate framing depth without confounding

Option A
Vary tone (warm vs neutral)
Make in-depth condition warmer or more supportive. Cleaner narrative but confounds disclosure depth with affective tone.
Option B
Vary explicit accuracy claims
Tell in-depth participants "the system is 95% accurate." Strong manipulation but confounds disclosure with credibility framing.
Option C
Vary mechanism disclosure only
Both texts use the same neutral tone. In-depth simply enumerates the four eye-behaviour metrics used to compute the score. No accuracy claims, no warmer language.

Why: Disclosure depth is what the literature points to (Nourani et al., 2019). Any other change (tone, accuracy claim) would mix manipulations and make the result impossible to attribute. The cost of isolation is a less dramatic manipulation. Worth it for interpretability.

Part III: Findings

Honest framing first. The a priori power calculation required N=190 per condition for the planned quiz contrast. Recruitment within the available lab time-frame, combined with pre-specified exclusion criteria (prior enrolment in the lecture's source course, incomplete eye-tracker logs), produced an analysis sample of N=15 (8 brief, 7 in-depth). All inferential results below are reported descriptively. The point is the pattern, not the confidence interval.

Headline results

EFFECT SIZE (COHEN'S D, IN-DEPTH VS BRIEF) 0 Perceived Accuracy +0.97 in-depth higher Perceived Intrusiveness +0.52 in-depth lower (small) Reactance (PRS-HCI) +2.00** in-depth lower Quiz score +0.01 · essentially zero
Framing reshapes how users feel about the system. It does not reshape what they learn from it. ** p < .01.

The in-depth framing condition produced substantially higher perceived system accuracy (d = -0.97, p = .086), much lower psychological reactance (d = +2.00, p = .008), and lower antipathy in particular (d = +1.78, p = .016). Quiz performance was indistinguishable between conditions (d = +0.01, p = .987).

Translation: how you introduce an AI feedback system at first interaction substantially changes how users feel about it. It does not, in this study, change how they perform on the underlying task.

The mediation model: only the first step replicates

The pre-specified serial mediation model traced the hypothesised pathway: framing → perceived accuracy → reactance → quiz performance. Bootstrap 95% confidence intervals from 5,000 resamples confirmed only the first path. Framing did shift perceived accuracy (+1.67, 95% CI [+0.38, +2.60]). The downstream path through reactance to quiz performance was directionally consistent (+1.95) but the confidence interval spanned nearly the full outcome scale. With N=15, the chain after the first step is indistinguishable from sampling noise.

Framing (X) +1.67* Perceived Accuracy -0.50 Reactance -2.32 Quiz (Y) — solid green: 95% bootstrap CI excludes zero — dashed grey: CI includes zero (not detectable at N=15)
First step holds. Downstream chain is directionally consistent but under-powered.

The eye-tracker triangulation: a second design lever

This is the finding that turned the case from "framing matters" into something more useful. Behavioural telemetry collected by the eye-tracker showed that participants in the in-depth condition received markedly fewer pop-ups during the lecture (M = 6.0) than those in the brief condition (M = 16.25). Across the whole sample, pop-up count correlated +.68 with reactance and -.70 with perceived accuracy.

That suggests an alternative pathway: in-depth framing may produce more sustained attention (because participants understood what was being measured), which led to fewer interrupt events, which led to lower reactance. The introduction does not just reshape perception. It reshapes the user's interaction with the system in a way that compounds.

For products that ship state-triggered feedback, this is the design implication that travels: frequency of interruption likely moves reactance more than the wording of the interruption does. Optimise the trigger logic before you optimise the message.

A confounder worth naming

Robustness analyses controlling for prior knowledge and topic interest attenuated the framing effect on perceived accuracy substantially (β = +0.56, p = .41), with topic interest emerging as a stronger predictor in its own right. With N=15 this cannot be cleanly separated, but it is a real signal: framing effects on perception may partially reflect baseline engagement differences, not just the manipulation. A confirmatory replication at the analytically required sample size would need to either randomise or measure-and-control for these covariates from the outset.

Design implications

Gerben Suerink
Gerben Suerink

Curious about the study?¿Te interesa el estudio?

The thesis goes deeper into the literature, the lab implementation, the full statistical analysis, and the discussion of the framing-vs-frequency tension for AI feedback products. Happy to share the manuscript, or talk through the design implications for products in this space.

La tesis profundiza en la literatura, la implementación de laboratorio, el análisis estadístico completo y la discusión sobre la tensión encuadre-vs-frecuencia para productos con feedback de IA. Encantado de compartir el manuscrito o de comentar las implicaciones de diseño para productos en este espacio.