◇ ◈ ◇
Agentic maths diagnosis · v0.5

Every learner has
hidden mines.

Most assessments find the wrong mine. They catch what a student got wrong, not why, and miss entirely what's underneath: the fixed mindset that quits at the first wall, the exam anxiety that blanks under pressure, the resilience (or absence of it) that shapes every result. Mindsweeper runs eight specialised AI agents across three assessment layers — Content (knowledge map), Control (mathematical character via validated psychometric instruments), and Capacity (cognitive profile via measured span tasks) — to sweep the complete terrain. What emerges is not a score. It's a diagnosis of how a student thinks, feels, and fights.

K–12·IB MYP/DP·IGCSE·A Level·O/IP Level·Gaokao·~30 minutes

How the sweep works

Four phases.
One complete picture.

Mindsweeper doesn't test. It probes. The way a great teacher would, if they had an hour alone with one student and the patience to truly look.

1
Know the terrain.
A short intake calibrates the diagnosis. Grade level, syllabus and goal set the question difficulty. Three validated self-report instruments (SRQ-A · MAI · CYRM-R Personal) map mathematical character. Digit-span and word-span tasks measure working-memory capacity. Then the student chooses a session mode.
Grade level
Syllabus
Goal
Portfolio upload
SRQ-A · Motivation
MAI · Metacognition
CYRM-R · Resilience
Digit span
Word span
→ Cognitive capacity
Open Socratic
Choose your question
Skip to report
2
Eight agents, three layers.
Evidence from any mode feeds a shared student model: the blackboard. Eight specialised AI agents read and write it simultaneously — five assessors across three layers (Content, Control, Capacity), one portfolio ingestion agent, one orchestrator, and one intervention planner. No finding is trusted from a single source; everything is cross-validated.
Socratic conversation
+
Past-work analysis
+
Question bank
+
Self-reports (SRQ-A · MAI · CYRM-R)
+
Span tasks
Shared student model
the eight agents
1
Orchestrator
Controller
Role
Reads all five assessors' open-gap lists each turn. Selects probe type (Reasoning, Perturbation, Teaching, Calibration, Frustration, Capacity, Structural). Assembles the final report: convergence analysis, verdict, and narrative. The system's attention mechanism.
Key principle
Prevents over-indexing on knowledge at the cost of character and capacity. If Control or Capacity has high uncertainty, a non-knowledge probe is injected regardless of Content's agenda.
2
Knowledge Mapper
Content layer · Assessor
Role
Classifies every response against the target curriculum strand (CCSS, IGCSE, IB, etc.). Tags as procedural or conceptual. Logs errors by type: careless slip, procedural bug, conceptual gap, or transfer failure. Root-gap analysis uses the SAP Coherence Map prerequisite graph.
Learning Commons · SAP Coherence Map · ANet decomposition
3
Mindset & Motivation
Control layer · Assessor
Role
Classifies motivation along Ryan & Deci's SDT autonomy continuum, validated by the SRQ-A (16-item shortened version). Separately tracks Dweck growth vs fixed mindset. Converges self-report scores with observed dialogue signals.
Instrument
SRQ-A — Self-Regulation Questionnaire (Academic), 16 items, stems A + D. Produces four subscale scores and a Relative Autonomy Index (RAI ±10).
4
Metacognition & Comms
Control layer · Assessor
Role
Compares predicted confidence with actual performance (calibration gap). Measures explanation quality via teach-back. Validated by the MAI (Harrison & Vallin 2018, 19-item shortened version). Converges self-report with in-session monitoring signals.
Instrument
MAI — Metacognitive Awareness Inventory, 19 items. Two subscales: Knowledge of Cognition + Regulation of Cognition.
5
Resilience & Affect
Control layer · Assessor
Role
Tracks behaviour under perturbation: cliff (sudden collapse) vs graceful slope. Flags exam anxiety markers. Validated by the CYRM-R Personal subscale (10 items, two-factor revised model). Converges self-report resilience with observed frustration tolerance.
Instrument
CYRM-R Personal — Child and Youth Resilience Measure (Revised), personal subscale, 10 items.
6
Cognitive
Capacity layer · Assessor
Role
Measures working-memory capacity via two controlled span tasks — adaptive digit span (up to length 10) and adaptive word span (sets of 2→5). Separately tracks ecological cognitive-load signals from the live session. Computes convergence between measured span and observed in-session load.
Key principle
Working-memory capacity cannot be read from Socratic dialogue — it requires controlled measurement. This distinction is architecturally enforced: span tasks slot in before the session, ecological signals come from within it.
Daneman & Carpenter 1980 · Conway et al. 2005 · CLT / Sweller
7
Portfolio Ingestion
Evidence adapter
Role
Converts uploaded PDFs/images into structured evidence objects. Extracts working steps, final answers, corrections and crossings-out, margin notes, and layout signals. A crossing-out is not an error — it may be self-correction, which is a resilience signal.
8
Intervention Planner
Consumer
Role
Reads the completed three-layer profile and generates a sequenced intervention plan. Knowledge sequencing via topological sort of the prerequisite graph. Mindset type gates the reframing strategy. Cites named science-of-learning principles for each action.
External tools
Calls the Learning Commons Knowledge Graph MCP for Content-layer sequencing: find_standard_statement, find_standards_progression, find_learning_components. Strategy library drawn from CLT, spaced retrieval (Cepeda), interleaving (Rohrer & Taylor), elaborative interrogation (Chi), and wise interventions (Yeager 2022).
Learning Commons · CLT · Roediger & Karpicke · Cepeda · Rohrer & Taylor · Chi · Dweck · Ryan & Deci
3
Clear the board.
Three outputs assembled by the Orchestrator across the three layers. The knowledge map (Content) names the root gap and its fan-out. The mathematical character profile (Control) via M³ — Mindset, Motivation, Metacognition — plus resilience, each converging self-report and observed signals. The cognitive profile (Capacity) from measured span tasks.
Content
Topic mastery by strand
Procedural vs conceptual
Error taxonomy
Root gap · fan-out
Control · M³
Motivation type (RAI)
Growth vs fixed mindset
Metacognitive awareness
Resilience profile
Self-report ↔ observed convergence
Capacity
Digit span
Word span
Ecological load signals
Span ↔ session convergence
4
Mark where to go next.
A targeted intervention plan grounded in the science of learning. Not generic advice. Four dimensions, each with specific actions. Different for every learner because the map is different for every learner.
Knowledge
Mindset
Habit
Problem-solving
View full technical architecture
Live demo

See it sweep.

A complete Mindsweeper session: from intake to finished report in 30 to 60 minutes.

Intake → Socratic probing → Portfolio analysis → Knowledge map → Character profile → Intervention plan

The report

Your map.
In full.

Three views (for teachers, parents, and students) from one unified profile. Diagnostic depth for the classroom; clear, actionable language for the home. Two dashboards available: the fully populated demo (Yu Xiao Ting sample case) and the fresh empty-state console ready for your first student.

Teacher view
Parent view
Student view
Demo dashboard → Fresh dashboard →

Demo: Yu Xiao Ting · Edexcel GCSE Foundation · Fresh: same structure, no data

Yu Xiao Ting · GCSE Foundation · v3
Course
GCSE Fdn
Edexcel 9-1
PT1 → PT2
71→56%
execution gap
Root gap
7.EE.A.2
fan-out = 4
Resilience
Mixed
2nd-guesses
P vs C
Recognised
unstable exec
Target
Perse Yr9
CEM · close gap first
Open demo dashboard →

Or view the fresh empty-state console · 7.EE.A.2 bottleneck · 4 critical topics · 8-week plan

What's next

The roadmap.

From working prototype to validated, persistent, multi-student platform.

Now · v0.5
8 agents · 3 layers · demo live
8-agent blackboard architecture
Content · Knowledge Mapper + curated question bank
Control · M³ via SRQ-A · MAI · CYRM-R
Capacity · Digit span + Word span tasks
Portfolio ingestion + Intervention planner
3-audience report · teacher console · feedback survey
11 curricula · standalone HTML deployment
Next · v1.0
Probe extensions · validation
Halpern structural-reasoning probes (G1–G7)
Adaptive probe sequencing
Validation study · 50+ students
EDF adaptive scaffolding
Reverse AI chatbot evaluation
Future · v2.0
Infrastructure · persistent memory
External memory · MongoDB
Persistent student records
Multi-session trajectory
MCP server deployment
Multi-student cohort view