Not All Tutoring Is The Same Tutoring: A Field Guide to Evaluating Modalities
Why it matters
“We have a tutoring program” tells you almost nothing anymore.
In-person, one-on-one instruction. A human tutor using an AI co-pilot. A chatbot a student logs into on their own. Districts increasingly lump all three under the same word — tutoring — and evaluate them with the same shrug of confidence.
They shouldn’t.
These are three different ways to instruct, with three different evidence bases, three different failure modes, and questions you need to ask before you buy.
This isn’t a ranking. It’s a map. Know which modality you’re looking at, and you’ll know which questions actually matter.
The 3 modalities
Before you can evaluate a tutoring product, you have to know what you’re actually evaluating. Broadly, what’s on the market today falls into three buckets:
In-Person High-Impact Tutoring (HIT) — human-delivered, in-person or live virtual, built on the design features research has repeatedly tied to strong outcomes.
AI-Augmented Human Tutoring (“hybrid” or “co-pilot”) — a human tutor still runs the session; AI supports them in the background with suggested prompts, feedback, and scaffolds.
AI-as-Tutor / Intelligent Tutoring Systems (ITS) — the student interacts primarily with software. No human tutor is present in the session itself.
“AI-powered tutoring” could mean either of the last two. Always ask which one you’re being sold.
6 questions to ask, modality by modality
1. Dosage. Is it happening or just scheduled?
The evidence-based minimum, per the National Student Support Accelerator (NSSA), is 3+ sessions/week, 30–60 min, a semester or longer. That’s the benchmark every modality should be measured against — not just on paper, but in practice.
- In-person HIT: dosage is the design. If it’s not happening at this frequency, it’s not HIT — it’s something weaker wearing HIT’s name.
- Hybrid/co-pilot: Can preserve HIT-level dosage if it’s still live tutoring — but Accelerate’s research on tutoring dosage warns districts to define “actual dosage,” not just scheduled dosage, and to keep tracking the gap between the two.
- AI-only/ITS: This is where dosage can get slippery fastest. Without school-day integration and active monitoring, “actual” usage is often far lower and more variable.
Ask: What’s the actual average weekly dosage per student, not the intended one — and how will the vendor produce that information?
2. Relationship: Is a consistent human still in it?
Same tutor, every session? Sustained tutor-student relationships are one of the defining, evidence-backed elements of high-impact tutoring — not a soft add-on. Annenberg’s EdResearch for Recovery brief names it a hallmark of effective programs.
- In-person HIT: Relationship is the mechanism. It’s built directly into the model.
- Hybrid/co-pilot: Relationship is retained — the tutor is still human — but NSSA notes real tradeoffs districts should watch between cost savings and personal connection as AI takes on more of the session.
- AI-only/ITS: The relationship element is reduced or absent. The evidence for comparable, durable impact without that relationship is much less established.
Ask: Does a student see the same person, session after session? If not, what is the model relying on instead — and does the evidence support that substitute?
3. Curriculum alignment: does it complement what’s happening in the classroom, or run parallel to it?
NSSA is clear that high-impact tutoring uses high-quality instructional materials (HQIM) and complements — not replaces or contradicts — classroom curriculum. For any modality, districts should require vendors to map to their HQIM and scope-and-sequence, or show evidence the skills being tutored are curriculum neutral.
- In-person HIT: Alignment is intentionally built into program design and easy to verify through tutoring materials and observing tutoring sessions.
- Hybrid/co-pilot: AI can help with planning and alignment prompts, but districts should still require the tool to map to their HQIM and scope-and-sequence — it won’t happen automatically.
- AI-only/ITS: Alignment risk is highest here. Many products aren’t tightly aligned to a specific district’s curriculum unless someone configures them to be, and that configuration is often skipped.
Ask: Can the vendor show, concretely, how this tool’s content maps to our curriculum — not just to “grade-level standards” in the abstract?
4. Tutor training & coaching: Who’s ensuring quality, and how would you know?
- In-person HIT: Tutors are trained and receive ongoing coaching through regular, periodic observations where fidelity is documented.
- Hybrid/co-pilot: AI support can actually help less-experienced tutors improve; NSSA has reported gains concentrated among students paired with lower-rated tutors when co-pilot tools were used. That’s a genuine strength of this modality — worth naming.
- AI-only/ITS: Training and coaching shift to “student onboarding” and product design. Quality assurance becomes far less transparent, and much harder for a district to actually monitor.
Ask: How will we know if quality is being maintained or not, and how quickly will we know?
5. Evidence strength: how much do we actually know?
- In-person HIT has a strong, multi-study evidence base. NSSA’s framework exists specifically because so much is known about what separates effective programs from ineffective ones.
- Hybrid/co-pilot is emerging but promising — NSSA points to randomized controlled trials showing AI embedded in live tutoring improved outcomes, including at least one large trial with effects replicated in a second context.
- AI-only/ITS evidence is mixed and uneven across products and settings. “AI tutoring” is not a single intervention with a single evidence base — and Annenberg’s research cautions that effects often shrink when programs scale without strong implementation.
Ask: Is the evidence for this specific product, or for the general idea of “AI in education”? Those are not the same thing. Ask vendors to provide the full research studies used for ESSA validation and any experimental research conducted by third parties.
6. Common failure modes: what usually goes wrong?
Every modality has a predictable way of failing. Knowing it in advance is half of prevention.
- In-person HIT fails through low dosage, inconsistent attendance, poor initial training, or thin coaching and fidelity monitoring.
- Hybrid/co-pilot fails when districts start treating the AI layer as a substitute for coaching and fidelity oversight, or let dosage quietly drift below the HIT threshold.
- AI-only/ITS fails through low genuine engagement, curriculum misalignment, “answer-giving” shortcuts that bypass actual learning, and weak district visibility into whether any of it is working.
Ask: Which of these failure modes are we already set up to catch — and which would we only discover a semester too late?
Putting It Together:
A Modality Isn’t Good or Bad. It’s a Set of Tradeoffs.
None of this is an argument that hybrid tools or AI-only products have no place. It’s an argument that each modality carries a different evidence base, a different relationship model, and a different set of things that can quietly go wrong — and treating them as interchangeable because they all get called “tutoring” is how districts end up disappointed.
This is also exactly why we’ve argued that no tutoring program — human, hybrid, or AI — should be added without a plan. Before adopting any modality, a district should still be able to answer the same questions we’ve raised before:
- What problem, for which students, are we actually solving?
- What does success look like, and by when?
- What are we willing to de-implement to make room for it?
- If it doesn’t work, how would we know — and what happens next?
The modality changes which risks you’re managing. It doesn’t change whether you have to manage them.
The Actual Bottom Line
“Tutoring” is not a single thing you can evaluate with a single checklist. In-person HIT, AI-augmented hybrid models, and AI-only tools each come with their own evidence base, their own relationship to the research on what works, and their own way of failing quietly.
Know which one you’re looking at. Ask the questions that modality actually demands. And hold every one of them — including MEC’s — to the same evidence-based standard.
Want help evaluating a specific tutoring product or modality for your district?
Let’s talk.
—Yours in service, Michigan Education Corps.
Sources:
National Student Support Accelerator — What Is High-Impact Tutoring?
https://nssa.stanford.edu/about/high-impact-tutoring
National Student Support Accelerator — District Playbook: Introduction to High-Impact Tutoring
https://nssa.stanford.edu/district-playbook/introduction-high-impact-tutoring
National Student Support Accelerator — Research Notes: Two Emerging Strategies Using AI in Tutoring
https://nssa.stanford.edu/news/research-notes-two-emerging-strategies-using-ai-tutoring
Accelerate — Defining Tutoring Dosage for Program Implementation and Applied Research
https://accelerate.us/research/defining-tutoring-dosage-for-program-implementation-and-applied-research/
Annenberg Institute / EdResearch for Recovery — Design Principles (Accelerating Student Learning with High-Dosage Tutoring)
https://annenberg.brown.edu/recovery/edresearch1




