PastAIED 2026 · 2026

Stakeholder-Driven Contextual Evaluation of Language Models in Education

27th International Conference on Artificial Intelligence in Education (AIED 2026)
Coex Center, Seoul, Republic of KoreaJune 27 – July 3, 2026

Theme: “From Tools To Teammates: Human-AI Synergy For Augmented Learning”

W14COEX 324June 27, 2026 (Saturday)1:00 – 6:00 PM (KST)

With the increasing reliance of AIED on opaque, black-box scaffolds such as large language models to support student learning, there is a growing concern about their limitations when used in diverse pedagogical contexts. This opacity often undermines stakeholders' trust and shapes their perceptions, contributing to resistance toward the adoption of AI scaffolds in schools.

To address these challenges, we developed AIBAT, a workflow and system designed to support stakeholders in auditing and critically evaluating the potential benefits and harms of AI systems within their specific pedagogical contexts (e.g., subject matter, grade level, English proficiency). With AIBAT, stakeholders can specify expected behaviors — i.e., what they anticipate the AI scaffold should do — and test the system against those expectations.

In this half-day tutorial, participants will use AIBAT to identify and make sense of AI-related risks and use evidence to calibrate their trust in AI scaffolds. At the end of the tutorial, we will deliberate on AI auditing processes and discuss broader implications for promoting responsible and effective stakeholder participation in the evaluation and deployment of AI systems in educational settings.

Topics

  • Stakeholder-driven evaluation of large language models
  • Contextual AI auditing in K–12 and higher education
  • Behavior analysis for equitable AI in classrooms
  • Responsible AI deployment in educational settings
  • Human–AI trust and transparency
  • Linguistic variation and fairness in AI grading models

Important dates

  • Tutorial DayJune 27, 2026 (Saturday)

Schedule

  1. 30 min

    Opening Remarks

    Discussion on the evaluation crisis with LLMs in educational settings.

  2. 30 min

    Small Group Deliberation

    Share-out on current approaches to AI evaluation across participants' contexts.

  3. 15 min

    Introduction to Behavior Analysis

    Overview of the behavior analysis framework underpinning AIBAT.

  4. 15 min

    Individual Think Time

    Participants identify relevant linguistic variations in their own use cases.

  5. 15 min

    Whole Group Demo of AIBAT

    Live walkthrough of AIBAT using a participant-contributed use case.

  6. 15 min

    Break

  7. 60 min

    Small Group Work with AIBAT

    Hands-on auditing session where participants evaluate an AI system using AIBAT.

  8. 30 min

    Fishbowl Demo

    Small group shares findings with the whole group in a fishbowl format.

  9. 30 min

    Closing Reflection

    Discussion on effective stakeholder participation in the evaluation and deployment of AI.

Related papers

  • AIBAT: AI Behavior Analysis Tool for Teacher-Driven Contextual Evaluation of Language Models in Education

    26th International Conference on Artificial Intelligence in Education (AIED) · 2025

    View paper
  • Evaluating Behaviors of General Purpose Language Models in a Pedagogical Context

    25th International Conference on Artificial Intelligence in Education (AIED) · 2024

    View paper
  • Identifying Biases in Large Language Model Assessment of Linguistically Diverse Texts

    Artificial Intelligence in Measurement and Education Conference (AIME-Con) · 2025

    View paper
  • Evaluating Language Models Using Linguistic Variations

    International Conference of the Learning Sciences (ICLS) · 2025

    View paper

Organizers

S
Shamya Karumbaiah

University of Wisconsin–Madison

A
Ananya Ganesh

University of Wisconsin–Madison

A
Anurag Maravi

University of Wisconsin–Madison