Math-first metacognitive tutoring · Research in progress

What if a math tutor noticed how you think?

ARIA is designed to support reflection while a student works through a math problem. It reads the reasoning rather than the final answer, then offers the smallest scaffold that could move the student forward. We are evaluating whether that improves persistence and clarity; current results are preliminary.

Prompts before hints, answers lastAnswers stay with the studentHuman validation is still pending
Reasoning-move exampleTap a move to change ARIA’s question
Student says

I know I have seen this kind of equation before, but I cannot tell which step comes next.

Observed: uncertainty + help-seeking
ARIA asks

Which step still feels solid? Start there, then name the first line where it stops making sense.

ARIA labels what is visible in the student’s words. It does not diagnose a hidden mental state, ability, emotion, or disability.

About ARIA

Built from personal experience with ADHD and research into today’s learning tools.We are studying how technology can strengthen independent thinking.

01

It started with lived experience.

Naren’s experience with ADHD made one problem clear: getting an answer is not the same as learning how to plan, work through confusion, and recover when a strategy fails.

02

Current tools often solve too much.

Our research into current tutoring tools found that many systems optimize for fast, correct responses. They rarely make the student’s thinking process visible or help students practice metacognition directly.

03

ARIA turns that gap into a research question.

Naren Saravanan and Karthik Malireddy are testing whether short, state-aware questions can help students plan and self-check independently. Success means the support becomes less necessary over time.

Why math first

ARIA is deliberately narrow.

Math is where a metacognitive tutor can be studied honestly before anyone claims it works elsewhere. The structure of a math problem is what makes the research question testable.

01

Problem-solving has clear steps.

A math problem breaks into discrete, observable moves — set up the equation, isolate the variable, check the result. That structure lets ARIA point to where a student stopped instead of guessing that they are generally confused.

02

Right and wrong are unambiguous.

Unlike an essay, a math step is either valid or it is not. That gives an honest correctness signal for whether a scaffold actually helped, and makes it harder for us to claim success without evidence.

03

Learning outcomes are measurable.

Attempts before giving up, time to a correct step, errors repeated across problems, and independent retries are all measurable in math. These are outcomes we can test rather than impressions we can only describe.

This is a scope choice, not a hidden limitation.

Starting with math means ARIA is not a general homework assistant, and we make no claims about subjects we have not studied. If the approach holds up here, extending it is a later research question rather than a feature we are shipping now.

How ARIA helps

Step-by-step scaffolding, not answer-giving.

The cycle stays on the step in front of the student. ARIA notices what they actually say, gives the smallest scaffold that could help, and looks for the support to become unnecessary.

  1. 01Notice

    The student works the problem out loud.

    A plan, a revision, a pause, or a question gives ARIA more to work with than a final answer alone.

    The student’s words stay central.
  2. 02Ground

    ARIA grounds itself in the step.

    It marks visible reasoning moves and keeps the exact evidence beside each label, at the level of a single step rather than the whole problem.

    The label describes the message, not the student.
  3. 03Scaffold

    ARIA offers the smallest useful scaffold.

    A prompt first, a narrowed hint only if that does not move the student, and a worked explanation only after repeated struggle.

    The student keeps ownership of the work.
  4. 04Transfer

    The support should fade.

    ARIA looks for the student to begin planning and checking without being prompted.

    Success means ARIA can step back.
Student says“Wait, I multiplied too early.”ARIA noticesSelf-correctionARIA asks“What will you check before trying again?”

The three-level help model

Help escalates only when the level before it did not work.

The student always receives the least amount of help that could still move them forward. An answer is the last resort, never the opening move.

  1. Level01

    Metacognitive prompt

    A question that returns attention to the student’s own reasoning. No mathematical content is given — the goal is for the student to state the plan or locate the breakdown themselves.

    What are you trying to find, and what is the first step you would try?
    Escalates if the student stays stuck or idle after responding
  2. Level02

    Guided hint or similar problem

    A narrowed hint about the specific step, or a simpler parallel problem that uses the same idea. The student still performs every step of their own problem; the hint only reduces the search space.

    Both sides still have an x. What could you do to get the x terms on one side?
    Escalates only after repeated attempts at this level
  3. Level03

    Answer explanation, after repeated struggle

    Only after documented struggle at the earlier levels does ARIA explain the reasoning for the step. The explanation focuses on why each move is made, and is followed by a similar problem the student completes independently.

    Here is why we subtract 3x from both sides — then try this one on your own.
    Never the first response, and never a silent solution

What ARIA watches for

Quiet does not always mean stuck.

The same behavior can mean very different things. A student who is quiet because the work is going well should not be interrupted the same way as a student who has quietly given up.

Stuck

Repeated attempts on the same step, backtracking, or no forward progress for an extended stretch.

ResponseA Level 1 prompt aimed at locating the exact step that broke down.

Confused

A step that contradicts the previous one, or a message that misstates what the problem is asking.

ResponseA question that checks understanding of the goal before any hint is given.

Idle

No input for a sustained period, with no partial work on screen — disengagement rather than thinking.

ResponseA low-pressure re-entry prompt, not a hint and not a nudge to hurry.

Struggling productively

Slow but real progress — the student is testing approaches, self-correcting, and moving between steps.

ResponseStay out of the way. Interrupting productive struggle is a failure mode, not a missed opportunity.

Rushing or misusing the system

Answers submitted faster than they could be worked, or repeated requests that skip straight to Level 3.

ResponseSlow the loop down and ask the student to show one step before help continues.

Planning

The student is setting up an approach before computing — the behavior we most want to strengthen.

ResponseAcknowledge and let it run. Prompting here is only useful if the plan is unstated.

Privacy-aware interaction

Text and lightweight choices first. Voice is optional.

Most interactions are a tap, a selected step, or one short line of text. Voice input is never required. Students work in classrooms, shared rooms, and beside peers, and asking someone to talk through a problem out loud can be exposing rather than helpful. Speaking should be a preference the student chooses, not a condition of getting help.

Built from student feedback

Moving from data we generated to students we listen to.

Research roadmap

Four phases, in order.

Each phase has to hold up before the next one is worth running, and we are early in that sequence. Naming where we are is more useful than implying we are further along.

  1. Phase 1Complete

    Synthetic think-aloud data

    Generate think-aloud examples to build and test a first version of the system. This shows the pipeline works end to end. It does not show that it works for real students.

  2. Phase 2Current

    Student interviews

    Structured interviews with students who struggle in math: where they get stuck, what help they have rejected, and how they want to interact. Findings replace our assumptions in the design.

  3. Phase 3Next

    Formative prototype testing

    Small sessions where students use the prototype on real math work while we observe. The goal is to find where prompts misfire, annoy, or get ignored, and revise before any outcome claims.

  4. Phase 4Planned

    Classroom and teacher-informed evaluation

    Evaluation designed with teachers, in real instructional settings, using measures they consider meaningful. Only at this stage would it be reasonable to discuss learning outcomes.

Where the work stands.

The foundation is in place: a task bank, a way to describe student reasoning, and a fair comparison study. The next step is to put each part in front of educators and students.

Task bank100

Problems ready for review

The bank includes answer guides, solution paths, common mistakes, hints, and scoring notes. Evaluation starts with the math tasks.

Built and checked in code · educator review is next
Student language13

Ways students show their thinking

ARIA can mark planning, checking, self-correction, uncertainty, and help-seeking while showing the exact words behind the label.

Working in the product · human annotation is next
Comparison study5

Versions of the system to compare

The same student moments will be tested with generic help, problem context, current reasoning, learning history, and the full ARIA pipeline.

Study designed · independent educator ratings are next
Classroom evidenceNot yet

Learning results

We have not run a classroom study, so we are not claiming that ARIA improves learning, retention, transfer, or ADHD outcomes.

A reviewed student study is still required
Development benchmark

Useful for finding failures. Not evidence of learning.

Synthetic examples · no human ground truth

Development checkCurrent resultPlain-language meaningEvidence level
Same-style synthetic recognition84.6%About 85 of 100 held-out simulated messages matched their designed label.Synthetic development test
Balanced synthetic score0.837Performance summarized while giving each legacy state equal weight.Synthetic development test
Writing-style stress test9.05-point gapAverage accuracy changed when unfamiliar generators wrote the examples.Cross-generator stress test
Independent human labelsPendingTwo trained annotators must label real student language before accuracy claims advance.No result yet
Why the old headline changed.

Synthetic labels can test software, but they cannot show that ARIA understands real students. The primary language target is now observable reasoning moves with exact evidence spans and independent human validation.

What happens next

Each stronger claim has a stronger evidence gate.

The protocol separates task correctness, response quality, language measurement, feasibility, learning, retention, and unprompted transfer.

  1. 01

    Have qualified educators independently review all 100 task models.

  2. 02

    Blindly rate five paired response conditions for grounding, actionability, learner ownership, and answer leakage.

  3. 03

    Validate observable reasoning moves on real student language, split by complete student or session.

  4. 04

    Run a reviewed feasibility pilot before testing learning and transfer against an active control.

Read the evaluation methodology

Limitations and current stage

What ARIA is not, yet.

ARIA is an early-stage research prototype. Stating that plainly is part of the work, so nothing on this site reads as more than it is.

  • The current model is built on synthetic think-aloud examples.

    The data was generated to imitate student reasoning. It was not collected from students in classrooms.

  • Reported results are development checks, not learning results.

    They show the software behaves as designed on examples we created. They say nothing about whether a student learned anything.

  • The interaction design is not yet validated.

    Prompt wording, escalation thresholds, and timing are informed judgments. Student interviews and prototype testing are how they get corrected.

  • We are not claiming classroom-validated learning gains.

    ARIA has not been shown to improve grades, retention, or transfer, and it is not a substitute for a teacher, a tutor, or an IEP support plan.

  • Real-student evaluation is the next step.

    The claims on this site will change as that evidence arrives, including if it contradicts what we expected.

What we are testing next

Four questions synthetic data cannot answer.

These are the questions student interviews and prototype sessions are designed to answer. Any of them could come back against us.

01

Do students find the prompts helpful?

A prompt a student experiences as nagging is worse than no prompt. We will ask students directly, in interviews and right after sessions, whether a given prompt helped, annoyed, or was ignored.

02

Does ARIA read student states accurately?

Detection is currently measured against generated text. The real test is agreement with what students report about their own experience, and with what a trained observer would say.

03

Does scaffolding improve persistence?

We are looking at whether students attempt more steps before asking for an answer, and whether they return to a hard problem — not at whether the final answer was right.

04

What input do students actually prefer?

Text, tappable choices, or voice, and under what conditions. We expect preference to depend heavily on setting, especially whether a student is working near peers.

Research inspiration

Ideas this work draws on.

ARIA is a student project. These are lines of education research that shaped how we think about the problem.

When a teacher steps in

Work on how experienced teachers decide when to intervene, and the recognition that help offered too early can cut short the struggle that produces learning. ARIA’s escalation model is an attempt to take that timing question seriously.

Engagement as changing states

Research describing learning as a sequence of states — engaged, confused, frustrated, disengaged — rather than one fixed measure of ability. This is why ARIA treats “stuck” and “struggling productively” as different situations.

Who gets attention

Studies of how teacher attention distributes unevenly across a classroom, and how quiet students can go unnoticed for long stretches. It motivates watching for idle and silent states, not only for students who ask for help.

These are influences on our thinking, not endorsements or affiliations. We are not affiliated with the researchers or institutions behind this work, and nothing here should be read as their validation of ARIA. A full reference list will accompany our paper.

October242026

Khan Lab School AI in Education Summit

Intentional Innovation: Keeping Learning Human in an AI World

Meet Naren Saravanan and Karthik Malireddy as they share ARIA’s research, current limitations, and next questions with educators, researchers, students, and builders.

When
Saturday, October 24 · 8:00 AM to 5:00 PM PT
Where
Khan Lab School · Mountain View, California
Get summit tickets

Founders

Built from lived experience. Tested with care.

ARIA began with a question shaped by experience with ADHD: what if a tutor paid attention to how a student was thinking instead of simply producing the next answer?

Student Researcher

Naren Saravanan

Senior, Marvin Ridge High School · Waxhaw, North Carolina

Lived experience, research direction, and the question at the center of ARIA.

Student Researcher

Karthik Malireddy

Senior, Marvin Ridge High School · Waxhaw, North Carolina

Co-research, system development, evaluation, and translating the idea into a testable tool.

ARIA needs more than a model. It needs people who know learning up close.

Contact the team
01

Teachers + research advisors

Interested in reviewing our study design? We would rather have it critiqued now than defend a flawed one later. Tell us which outcomes are worth measuring, and help us evaluate ARIA with real students.

Review our study design
02

Students

Have you struggled with math tutoring tools? Share your experience: where you get stuck, which tools you stopped using, and whether you would type, tap, or speak. You do not need to be good at math to help.

Share your experience
03

Researchers + developers

Help with real think-aloud datasets, human annotation, study design, or interface work for low-effort, privacy-aware input.

Start a conversation
Research updates

Follow the honest version of the story.

New evidence, limitations, demos, and ways to participate, sent only when there is something useful to share.

Join research updates