Research note 01 · Updated July 2026
How ARIA plans to earn its claims.
ARIA separates software checks, educator judgment, real-language validation, student feasibility, and learning outcomes. A result advances only the claim it actually tests.
The research question
Can ARIA provide grounded help without taking over the work?
The program tests several different questions: whether tasks are correct, whether responses use the student’s actual reasoning, whether language labels match independent human judgments, and eventually whether students learn and transfer strategies.
Evidence sequence
Five gates, from a working system to a learning claim.
- 01
Validate the task bank.
Two subject-qualified educators independently review answer models, solution paths, misconceptions, hints, and scoring criteria. Disagreements stay visible.
- 02
Blindly compare response quality.
Educators rate five paired conditions for problem grounding, student grounding, usefulness, actionability, ownership, answer leakage, and invented student actions.
- 03
Validate language on real student messages.
Two trained annotators label observable reasoning moves with exact evidence spans. Evaluation is split by complete student or session, never random messages.
- 04
Run a reviewed feasibility pilot.
With institutional review, parent permission, and student assent, test whether the tool is understandable, usable, and safe before asking whether it is effective.
- 05
Measure learning against an active control.
Use independent outcome tasks, concealed assignment, blinded scoring, intention-to- treat analysis, and a delayed no-ARIA transfer measure.
Current evidence ledger
Structured task drafts
Schema-valid and ready for independent educator review.
Observable reasoning moves
Each automatic label keeps the exact supporting words visible.
Completed outcome studies
No classroom learning or ADHD-specific effectiveness claim is being made.
| Evidence layer | Current status | What remains | Claim allowed today |
|---|---|---|---|
| Task models | 100 schema-checked drafts | Educator correctness review is pending. | The bank is structured and ready for review. |
| Observable language | 13 reasoning-move codes | Independent annotation on real student language is pending. | ARIA can show the exact phrase behind a tentative label. |
| Response quality | 5 blinded conditions designed | Two qualified educators must rate the locked responses. | A fair comparison can be run; no winner is claimed yet. |
| Synthetic stress test | 84.6% same-style accuracy | The balanced score is 0.837; the cross-generator gap is 9.05 points. | Useful for debugging only. |
| Student outcomes | 0 completed controlled studies | Feasibility, learning, retention, and transfer remain untested with students. | No effectiveness claim. |
Reading the language
Four terms that prevent inflated claims.
- Development benchmark
- A test used to find software weaknesses. Synthetic benchmark scores are not evidence that students learn more.
- Human ground truth
- Labels created independently by trained people using a locked codebook, without seeing ARIA’s prediction.
- Active control
- A comparison tool with the same tasks, interface, time, and base model, but without ARIA’s learner-conditioned pipeline.
- Transfer
- A student uses planning, checking, or self-correction on a new task while ARIA is absent.
Sources behind the design
Evidence for the method, not proof of the product.
These sources support how ARIA should be studied. They do not establish that ARIA is effective.
- Observable self-regulated learning
- Human coding of planning, monitoring, evaluation, and related think-aloud activity.
- Human evaluation of AI tutors
- Mistake identification, guidance, and actionability evaluated with human labels.
- Grounded tutoring dialogue
- Teacher-authored scaffolding and documented risks of incorrect feedback or answer revelation.
- Metacognition guidance
- Planning, monitoring, and evaluation taught explicitly inside subject learning.
- What Works Clearinghouse standards
- Randomization, attrition, baseline equivalence, eligible outcomes, and study confounds.
- Research with children
- Institutional review, parental permission, affirmative assent, and risk requirements.
Limitations
What ARIA does not know yet.
- The 100 task models have not yet been approved by independent educators.
- The observable-move baseline has not yet been tested against real human labels.
- The five response conditions have not yet received locked, blinded ratings.
- No student study has established usability, learning, retention, or transfer.
- Synthetic benchmark performance does not establish real-student understanding.
- Personal experience with ADHD motivates the question; it is not clinical evidence.
Evidence base
Built from learning science and rigorous evaluation standards.
The design draws on observable self-regulated-learning coding, human evaluation of AI tutoring, grounded tutoring dialogue, metacognitive transfer research, and What Works Clearinghouse standards. Those sources justify the study design, not ARIA’s effectiveness.