Why AI answers everything but never teaches you to ask
The EduNext 2026 report by the Look4ward Observatory (Luiss and Intesa Sanpaolo) carries a line in a pull quote on page 41 that says nothing about models, platforms or budgets. It is about posture, which means the way you position yourself in front of the machine before you even type the first word.
«Use these tools in a Socratic way, not to find answers but to find questions. In an age when answers are always available, the rare skill is knowing which question to ask, and that is not something you learn by delegating to the machine.»
Read quickly it sounds like plain common sense, the kind of sentence everyone nods at and nobody acts on. The trouble is that here common sense runs against instinct, because a tool that answers anything in three seconds trains you to request answers rather than to build questions. And thirty years of research say two uncomfortable things. The first is that asking does not come naturally to us even without AI in the room. The second is that precisely because of this, the Socratic mode is not a character trait you either have or lack, but a skill that can be trained and measured.
Asking does not come naturally

Who actually brings the questions into a classroom. Source Graesser and Person, 1994.
In 1994 Arthur Graesser and Natalie Person did something simple and rarely done, which was counting. Not estimating, counting, recording questions in real classrooms and in one-to-one tutoring sessions. A teacher asks 69 questions an hour on average, with peaks reaching 120. A whole class, every student added together, reaches a median of 3 questions an hour. A single student, the one sitting in front of you, asks 0,11 an hour, which is roughly eleven questions across a hundred hours of teaching.
That same student, placed in one-to-one tutoring, asks around 240 times more. Nobody became more curious in a week, the context changed. Curiosity is not the bottleneck. The bottleneck is the situation in which a question becomes possible, sensible and socially costless.
There is a second layer, and it is the one that really matters. Not all questions do the same work. In the study's two samples, deep-reasoning questions, the ones that ask for a why, a mechanism or a consequence, made up 22% and 39% of the total, and classifying everything with Bloom's taxonomy only 8% of student questions went beyond the first level. The decisive finding comes at the end, since the frequency of questions was not correlated with exam results, while the share of deep questions was, with a correlation of 0,47. How much you ask does not matter. What kind of question you ask does.
Change the question and the result changes

The measured effects of interventions that teach question building. Source Rosenshine, Meister and Chapman, 1996.
If asking well is not spontaneous, the next question is whether it can be taught. The answer has been available for thirty years and is still barely used in schools. Barak Rosenshine, Carla Meister and Saul Chapman reviewed 26 intervention studies in which students were taught to generate questions about the text they were reading. The median effect size on comprehension is 0,86 on researcher-built tests and 0,36 on standardised ones. In educational research an effect above 0,80 counts as large, and the highest value in the whole review, 1,12, comes from the most mundane support imaginable, which is handing students a few ready-made question openings to build their own on.
This finding has a direct counterpart in AI use. A 2026 study published in the International Journal of Educational Technology in Higher Education logged every single message that twenty-two postgraduate students wrote to a conversational assistant during a sixty-minute task. The researchers measured the share of messages that sought an explanation, meaning those containing a why, a how, an explain, a justify. Holding the number of exchanges constant, a one standard deviation difference in that share was worth roughly six extra points out of a hundred in the quality of the work delivered. It did not change immediate recall of the content, but it changed the quality of the output. The sample is small and the signal should be taken for what it is, namely something consistent with the earlier literature rather than conclusive proof.
A Socratic tutor is not enough on its own

The most frequent requests to a tutor built to ask questions. Source Hashmi and Rebello, 2026.
Here comes the part that should concern anyone about to buy a tool and consider the problem solved. In August 2026 two Purdue University researchers classified 2.874 messages written by 240 university physics students to an artificial tutor explicitly designed to question rather than to solve. The first twenty-five categories cover about 52% of all messages, so the behaviour is far more concentrated and repetitive than you would expect.
The most frequent move is writing the energy equation, at 8,3%. The second, at 4,4%, is asking the tutor what the next step is. The authors call it a meta-procedural request and read it as handing over command, because the student stops deciding the strategy and asks the machine to decide instead. Across the whole ranking only one request asks for a mechanism, namely the physical principle at work, and it accounts for 1,5%. The tutor was Socratic. The students, for the most part, were not.
The same pattern shows up in working adults. A Microsoft Research and Carnegie Mellon study presented at the CHI 2025 conference collected 936 real cases of AI use reported by 319 professionals. The more they trusted the tool, the less critical thinking they reported exercising. The more they trusted their own competence on the task, the more they exercised. And critical thinking does not disappear, it shifts, moving away from gathering information and framing the problem towards verifying output and supervising the process. Put differently, the first thing we delegate is exactly the part where you decide what question you are asking. We have already covered the cost of that delegation when we wrote about the AI learning penalty.
Two hours of training change how you ask

The effect of a two-hour workshop on questioning behaviour. Source Clerc and colleagues, 2026.
The good news is that posture moves, and with a surprisingly small investment. In 2026 a French research group worked with 116 students aged 13 to 15, splitting them into 76 trained and 40 controls. The training was a two-hour workshop, held two days before the test, explaining how a language model works, where it fails, how a request is built and how a response is evaluated. Everyone then tackled six science investigation tasks with a chatbot available.
Faced with a badly framed request, trained students asked for clarification in 59,2% of cases against 27,9% of untrained ones. They accepted an underspecified request without rewriting it in 51,5% of cases against 66,7%. The final score was 11,4 out of 20 against 10,3, a statistically significant difference. One detail is worth more than all the others, since when the students were asked to rate their own AI skills and their own metacognitive awareness, the answers bore no relation to actual performance. Feeling competent has nothing to do with being competent, and that holds for students as much as for adults.
Dosage matters too. In November 2025 Google's LearnLM team, together with the Eedi platform, ran a controlled trial with 165 students aged 13 and 14 across five UK schools, comparing static hints, human tutors and an artificial tutor whose every message was approved by a human tutor before reaching the student. On immediate mistake remediation the supervised system scored 93,0% against 91,2% for human tutors and 65,4% for static hints. On transfer to the next topic, which is the measure that really counts, it scored 66,2% against 60,7% for human tutors. Tutors edited 25,6% of messages, and 44,3% of their edits served to slow things down, because the tutor kept questioning beyond the student's patience. The Socratic mode works, but it has a dose beyond which it becomes an irritation and learning shuts down.
The ladder of the question

Five ways of addressing the same machine. Source Retoria elaboration.
I bring this ladder into the classroom, because I find it more useful than any list of best practices. These are five ways of addressing the same tool, and they completely change what is left in you when you close the laptop.
The first rung is delegation, meaning «just do the task for me». You get a result and learn nothing, and this is the case where research measures the damage.
The second is execution, meaning «what is the next step». It looks like a question, but it is the request the Purdue students made most often, and it actually hands the machine the valuable part of the work, which is deciding where to go.
The third is mechanism, meaning «explain why it works this way». Here you start genuinely asking, and these are the requests that in the 2026 study were worth six extra points on the final work.
The fourth is checking, meaning «where could this answer be wrong». This is the rung that flips the relationship of trust, because you stop treating the output as a verdict and start treating it as a hypothesis. The Microsoft and Carnegie Mellon research says this is exactly where the critical thinking of people who work well with these tools concentrates.
The fifth is reframing, meaning «what question did I fail to ask». This is the Socratic rung in the full sense, since you use the machine to discover the perimeter of your own ignorance rather than to fill a gap you have already identified. It is the move the EduNext report describes as the rare skill, and it is also the only one no tool can make on your behalf, given that it grows out of the suspicion that you framed the problem badly.
The two rules I always give are trivial and they work. First, before writing to the machine write one line about what you want to understand, not what you want to obtain. Second, when the answer arrives make one extra move, which is asking where it could be wrong. They cost less than a minute and they move the conversation up two rungs.
What this means for training
There is a strong temptation right now, which is to think the problem is the tool and the solution is picking the right one. The data say otherwise, namely that the same tool produces opposite outcomes depending on who uses it and how they were prepared to use it. The Purdue Socratic tutor was well designed and students used it as an executor. The generic chatbot in the French study was the same for everyone, and two hours of training changed the behaviour of those who had them.
This is why at Retoria we do not teach tools, which expire, we train posture, which stays. Teaching people to ask is a piece of educational work as old as Socrates and measured by research for thirty years, and AI has suddenly made it the most profitable skill a school or an organisation can develop.
If you want to know where your school or organisation stands, write to us for an AI readiness assessment, look at how we work, or read how value is shifting from memory to judgment.
Methodological note
The opening quotation comes from page 41 of the EduNext 2026 report, chapter 2, and is the anonymised voice of one of the thirty-two executives interviewed between December 2025 and April 2026. Its punctuation has been adapted without altering its content, as with the other quotations from the report that we publish.
The Graesser and Person (1994) figures come from two tutoring samples, one on a research methods course and one on an algebra course, and the value of 0,11 questions per hour for a single student is the study's estimate based on classroom observation. The factor of roughly 240 times is the ratio computed by the authors. The correlation of 0,47 between deep questions and exam outcomes refers to the second half of the course.
The effect sizes from Rosenshine, Meister and Chapman (1996) are medians across 26 intervention studies and not the effects of a single experiment. The 2026 study on interaction with language models has a sample of twenty-two participants in a single-group design, so it should be read as an indication rather than conclusive evidence. The Microsoft Research and Carnegie Mellon study is based on self-reported cases, so it measures perceived rather than objectively assessed critical thinking. The Purdue taxonomy is an August 2026 preprint and the percentages indicate the frequency of message categories, not of students. The LearnLM results concern an artificial tutor under constant human supervision, not a system left free to answer.
This article is not individual pedagogical advice, nor a clinical assessment of cognitive ability.
Sources
- Look4ward Observatory, Luiss Research Center for Strategic Change "Franco Fontana" & Intesa Sanpaolo (2026). EDUNext: Nuovi scenari per l'Education e le competenze nell'era dell'IA (CC BY 4.0 licence).
- Graesser, A. C. & Person, N. K. (1994). Question Asking During Tutoring. American Educational Research Journal, 31(1), 104-137.
- Rosenshine, B., Meister, C. & Chapman, S. (1996). Teaching Students to Generate Questions. A Review of the Intervention Studies. Review of Educational Research, 66(2), 181-221.
- Tsiligkiris, V. (2026). What students ask matters. LLM interaction depth, task quality, and immediate recall in higher education. International Journal of Educational Technology in Higher Education, 23(1), 44.
- Hashmi, S. F. A. & Rebello, N. S. (2026). A bottom-up taxonomy of student discourse with a Socratic AI physics tutor. Purdue University, preprint.
- Lee, H. P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R. & Wilson, N. (2025). The Impact of Generative AI on Critical Thinking. Proceedings of CHI 2025.
- Clerc, O., Abdelghani, R., Desvaux, C., Poisson, E., Oudeyer, P. Y. & Sauzéon, H. (2026). Teaching Students to Question the Machine. Preprint.
- LearnLM Team, Google & Eedi (2025). AI tutoring can safely and effectively support students. An exploratory RCT in UK classrooms.