Independent research in metacognition

What if a learning tool noticed how you think?

ARIA studies a student’s reasoning, not just the final answer, then asks one useful question to help them plan, recover, or check their work.

Evidence spans stay visibleAnswers stay with the studentHuman validation is still pending
Reasoning-move exampleTap a move to change ARIA’s question
Student says

I know I have seen this before, but I cannot tell which step comes next.

Observed: uncertainty + help-seeking
ARIA asks

What part still feels clear? Start there, then name the first point where the path gets fuzzy.

ARIA labels what is visible in the student’s words. It does not diagnose a hidden mental state, ability, emotion, or disability.

About ARIA

Built from personal experience with ADHD and research into today’s learning tools.We are studying how technology can strengthen independent thinking.

01

It started with lived experience.

Naren’s experience with ADHD made one problem clear: getting an answer is not the same as learning how to plan, work through confusion, and recover when a strategy fails.

02

Current tools often solve too much.

Our research into current tutoring tools found that many systems optimize for fast, correct responses. They rarely make the student’s thinking process visible or help students practice metacognition directly.

03

ARIA turns that gap into a research question.

Naren Saravanan and Karthick Malireddy are testing whether short, state-aware questions can help students plan and self-check independently. Success means the support becomes less necessary over time.

Founders

Built from lived experience. Tested with care.

ARIA began with a question shaped by experience with ADHD: what if a tutor paid attention to how a student was thinking instead of simply producing the next answer?

Student Researcher

Naren Saravanan

Senior, Marvin Ridge High School · Waxhaw, North Carolina

Lived experience, research direction, and the question at the center of ARIA.

Student Researcher

Karthick Malireddy

Senior, Marvin Ridge High School · Waxhaw, North Carolina

Co-research, system development, evaluation, and translating the idea into a testable tool.

How ARIA works

Four moments. One direction: more independence.

ARIA’s workflow is sequential on purpose. Each intervention begins with evidence and ends by returning control to the learner.

Notice

The student thinks out loud.

Words, revisions, pauses, and typing rhythm reveal more than a final answer can. ARIA pays attention to the learning process while the student works.

The work stays on the device.
S

I think I multiply first… wait.

Ground

ARIA identifies observable reasoning moves.

The system marks visible moves such as planning, justification, checking, self-correction, uncertainty, and help-seeking, then preserves the exact words supporting each label.

The label describes the message, not the student.
Observed moveSelf-correctionEvidence stays visible
  • “Wait”
  • “I meant subtract”
  • No diagnosis inferred
Ask

One question interrupts the pattern.

ARIA does not hand over the solution. It chooses a short Socratic prompt that helps the student plan, check, recover, or reflect.

The student keeps ownership of the work.
ARIA asks

“Which part of your plan still feels reliable?”

No answer revealed
Transfer

The prompt should become unnecessary.

Over time, ARIA looks for the student to begin planning and self-checking independently. That transfer, not more time with an AI, is the long-term research goal.

Success means ARIA can step back.
ARIA prompts
Independent planning
Step back as the student steps forward.

Research infrastructure first. Human evidence next.

ARIA now has a testable research program, not just a model demo. These numbers describe what is built and what remains unproven.

Research asset 01100

Structured math and English tasks

Every research task now includes acceptable answers, solution paths, misconception evidence, graded hints, scoring criteria, and provenance.

Schema checks pass · independent educator review pending
Research asset 0213

Observable reasoning moves

ARIA records visible moves such as planning, justification, checking, self-correction, uncertainty, and help-seeking with the exact words supporting each label.

Transparent baseline · independent human validation pending
Research asset 035

Blinded evaluation conditions

The locked study compares generic, problem-only, turn-grounded, profile-and-history, and full closed-loop responses on the same tasks.

100 paired episodes planned · two qualified educators required
Evidence boundary0

Completed classroom outcome studies

ARIA has not yet shown that it improves learning, retention, transfer, or outcomes for students with ADHD. Those claims require reviewed studies with real students.

Important limitation · causal evidence remains pending
Development benchmark

Useful for finding failures. Not evidence of learning.

Synthetic examples · no human ground truth

Development checkCurrent resultPlain-language meaningEvidence level
Same-style synthetic recognition84.6%About 85 of 100 held-out simulated messages matched their designed label.Synthetic development test
Balanced synthetic score0.837Performance summarized while giving each legacy state equal weight.Synthetic development test
Writing-style stress test9.05-point gapAverage accuracy changed when unfamiliar generators wrote the examples.Cross-generator stress test
Independent human labelsPendingTwo trained annotators must label real student language before accuracy claims advance.No result yet
Why the old headline changed.

Synthetic labels can test software, but they cannot show that ARIA understands real students. The primary language target is now observable reasoning moves with exact evidence spans and independent human validation.

What happens next

Each stronger claim has a stronger evidence gate.

The protocol separates task correctness, response quality, language measurement, feasibility, learning, retention, and unprompted transfer.

  1. 01

    Have qualified educators independently review all 100 task models.

  2. 02

    Blindly rate five paired response conditions for grounding, actionability, learner ownership, and answer leakage.

  3. 03

    Validate observable reasoning moves on real student language, split by complete student or session.

  4. 04

    Run a reviewed feasibility pilot before testing learning and transfer against an active control.

Read the evaluation methodology
October242026

Khan Lab School AI in Education Summit

Intentional Innovation: Keeping Learning Human in an AI World

Meet Naren Saravanan and Karthick Malireddy as they share ARIA’s research, current limitations, and next questions with educators, researchers, students, and builders.

When
Saturday, October 24 · 8:00 AM to 5:00 PM PT
Where
Khan Lab School · Mountain View, California
Get summit tickets

ARIA needs more than a model. It needs people who know learning up close.

Contact the team
01

Researchers

Help with real think-aloud datasets, human annotation, study design, or new cognitive-state taxonomies.

Start a conversation
02

Educators

Share what students with ADHD, dyslexia, and other learning disabilities need from a responsible pilot.

Start a conversation
03

Families + students

Tell us what feels supportive, what feels intrusive, and what an AI tutor should never do.

Start a conversation
Research updates

Follow the honest version of the story.

New evidence, limitations, demos, and ways to participate, sent only when there is something useful to share.

Join research updates