Hanademi

When AI does the homework

28 slides · 17 min · 2026-07-11 language
ENES
theme
LightDark
view
DeckTableTalk
brand
HanademiPlatzi
When AI does the homeworkMade for Freddy Vega, by Hanademi
Cuando la IA hace la tareaMade for Freddy Vega, by Hanademi
ChatGPT clarifies more than it buildscritical thinkingUniversity student responses across 120 countries, percentMade for Freddy Vega, by HanademiSources: NoteGPT. (2026, May 6). What 22,963 students told us about AI for homework. Analysis of Mendeley Data DOI10.17632/nv2343nwsb.2.Unit% of respondents73%Clarifying54.9%AI literacy43%Reliable39.1%Critical thinking
Clarity led at 73%, while critical thinking ranked last at 39.1%.
ChatGPT aclara más de lo que fomenta elpensamiento críticoRespuestas de universitarios en 120 países, porcentajeMade for Freddy Vega, by HanademiFuentes: NoteGPT. (2026, May 6). What 22,963 students told us about AI for homework. Analysis of Mendeley Data DOI10.17632/nv2343nwsb.2.Unidad% de encuestados73 %Aclaratorio54,9 %Alfabetización en IA43 %Confiable39,1 %Pensamiento crítico
La claridad lideró con 73%, mientras el pensamiento crítico quedó último con 39.1%.
Homework rose as independent exams fellChange after generative AI adoption among 26,811 Chinese students in grades 7-12, percent.Made for Freddy Vega, by HanademiSources: Strömberg, D., Lei, V., & Wu, Y. (2026). The generative AI learning penalty: Evidence from Chinese secondary education.CEPR Discussion Paper 21577.Hanademi analysis: Across more than five hundred thousand grades, the A-grade share rose thirteen points overall and sixteen in homework-heavy courses,while a separate study of twenty-six thousand eight hundred eleven students found homework scores up eighteen percent but entrance-exam scores down as…Unitpercent changeHomework18%Monthly exam-20%Entrance exam low-18%Entrance exam high-24%
Start with the central split. Homework scores rose 18%, but monthly exams fell 20% and entrance exams fell as much as 24%. AI can look like an exoskeleton during practice while the exam still measures the muscles underneath.
Las tareas subieron mientras los exámenesindependientes cayeronCambio tras la adopción de IA generativa entre 26.811 estudiantes chinos de 7.º a 12.º grado, porcentaje.Made for Freddy Vega, by HanademiFuentes: Strömberg, D., Lei, V., & Wu, Y. (2026). The generative AI learning penalty: Evidence from Chinese secondary education.CEPR Discussion Paper 21577.Análisis de Hanademi: En más de quinientas mil calificaciones, la proporción de sobres subió trece puntos en general y dieciséis en cursos con muchastareas, mientras otro estudio de veintiséis mil ochocientos once estudiantes halló que las tareas subieron dieciocho por ciento pero los exámenes de…Unidadcambio porcentualTareas18 %Examen mensual-20 %Ingreso, estimación baja-18 %Ingreso, estimación alta-24 %
Empecemos con la división central. Las notas de tareas subieron 18%, pero los exámenes mensuales cayeron 20% y los de ingreso hasta 24%. La IA puede parecer un exoesqueleto durante la práctica, mientras el examen sigue midiendo los músculos que hay debajo.
Brown’s suspected AI case produced astark assessment reversalScores in one Brown University course, including the historical final-exam range; AI use and cheatingremain allegations.Made for Freddy Vega, by HanademiSources: The Register. (2026, July 9). Brown says AI make class dumb, teacher must help use better.; Inside Higher Ed. (2026, July8). Brown professor suspects majority of his class used AI to cheat.Hanademi analysis: At least fifty of eighty-six students is fifty-eight point one percent, while ninety-six minus forty-eight point six is a forty-seven pointfour-point collapse, equal to forty-nine point four percent of the midterm score, and the final was sixteen point four points below the historical low of sixty-five.Unitpercent scoreTake-home midterm96%Controlled final48.6%Historical low65%Historical high80%
This is a single course, and the AI attribution has not been finally adjudicated. Even so, the assessment contrast is extraordinary: 96% at home, then 48.6% under control. The episode raises a question larger than cheating, whether the original assessment measured student knowledge at all.
El presunto caso de IA en Brown produjouna fuerte reversiónPuntuaciones de un curso de Brown University, incluido el rango histórico del examen final; el uso de IA y elfraude siguen siendo acusaciones.Made for Freddy Vega, by HanademiFuentes: The Register. (2026, July 9). Brown says AI make class dumb, teacher must help use better.; Inside Higher Ed.(2026, July 8). Brown professor suspects majority of his class used AI to cheat.Análisis de Hanademi: Al menos cincuenta de ochenta y seis estudiantes equivale a cincuenta y ocho coma uno por ciento, mientras noventa y seismenos cuarenta y ocho coma seis produce una caída de cuarenta y siete coma cuatro puntos, equivalente al cuarenta y nueve coma cuatro por…Unidadpuntuación porcentualParcial en casa96 %Final supervisado48,6 %Mínimo histórico65 %Máximo histórico80 %
Es un solo curso y la atribución a la IA no ha sido resuelta de forma definitiva. Aun así, el contraste es extraordinario: 96% en casa y después 48,6% bajo supervisión. El episodio plantea una pregunta mayor que el fraude: si la evaluación original medía realmente el conocimiento del alumno.
One suspected case became a warningabout assessment trustA Brown professor alleged that at least 50 of 86 students cheated on one midterm, while theuniversity had not issued final findings.Sources: Pascual, M. G. (2026, June 28). Professor denounces mass AI fraud on an exam at Brown University. El País.
Behind the scores is a professor confronting a basic loss of trust. He said his evidence implicated at least 50 students in a class of 86, but the university had not completed its findings. That distinction matters because suspicion is not adjudication.
Un caso sospechoso se convirtió en unaalerta sobre la confianzaUn profesor de Brown afirmó que al menos 50 de 86 estudiantes hicieron fraude en un parcial,mientras la universidad aún no había emitido conclusiones finales.Fuentes: Pascual, M. G. (2026, June 28). Professor denounces mass AI fraud on an exam at Brown University. El País.
Detrás de las puntuaciones hay un profesor ante una pérdida básica de confianza. Afirmó que sus pruebas implicaban al menos a 50 estudiantes de una clase de 86, pero la universidad no había concluido su investigación. Esa distinción importa porque una sospecha no es una resolución.
AI became normal assessmentinfrastructure in one yearUK undergraduates reporting generative AI use for assessments, 2024 and 2025.Made for Freddy Vega, by HanademiSources: Freeman, J. (2025). Student Generative AI Survey 2025. Higher Education Policy Institute and Kortext.Unitpercent of students202453%202588%1.7x
In 2024, just over half of UK undergraduates used generative AI for assessments. One year later, the figure was 88%. This measures use, not cheating, but it makes prohibition alone an implausible operating model.
La IA se volvió infraestructura habitual deevaluación en un añoUniversitarios británicos que declararon usar IA generativa en evaluaciones, 2024 y 2025.Made for Freddy Vega, by HanademiFuentes: Freeman, J. (2025). Student Generative AI Survey 2025. Higher Education Policy Institute and Kortext.Unidadporcentaje de estudiantes202453 %202588 %1,7x
En 2024, poco más de la mitad de los universitarios británicos usaba IA generativa en evaluaciones. Un año después, la cifra era 88%. Esto mide uso, no fraude, pero convierte la prohibición por sí sola en un modelo operativo poco plausible.
AI already reaches 84% of high schoolstudentsHigh school students reporting generative AI use for schoolwork, May 2025.Made for Freddy Vega, by HanademiSources: College Board. (2025, May). New research: Majority of high school students use generative AI for schoolwork.Unitpercent of students84%Schoolwork use
By May 2025, 84% of high school students reported using generative AI for schoolwork. That includes legitimate assistance as well as possible substitution. At this penetration, AI literacy and independent verification become basic academic skills.
La IA ya llega al 84% de los estudiantes desecundariaEstudiantes de secundaria que declararon usar IA generativa para tareas escolares, mayo de 2025.Made for Freddy Vega, by HanademiFuentes: College Board. (2025, May). New research: Majority of high school students use generative AI for schoolwork.Unidadporcentaje de estudiantes84 %Uso para tareas escolares
Para mayo de 2025, el 84% de los estudiantes de secundaria declaraba usar IA generativa para tareas escolares. La cifra incluye ayuda legítima y posible sustitución del trabajo propio. Con esta penetración, el dominio de la IA y la verificación independiente se vuelven habilidades académicas básicas.
A grades rose most where work waseasiest to outsourceChange in A-grade share after ChatGPT across more than 500,000 university grades, percentage points.Made for Freddy Vega, by HanademiSources: Chirikov, I. (2026). Study of AI exposure and more than 500,000 university grades, summarized by The Decoder.Unitpercentage points16%Homework-heavy courses13%All courses
Across more than half a million grades, the share of A grades rose 13 points after ChatGPT. In homework-heavy courses, the increase reached 16 points. Better grades are visible, but this design does not prove that underlying knowledge improved.
Los sobresalientes subieron más donde eramás fácil delegarCambio en la proporción de sobresalientes tras ChatGPT en más de 500.000 calificaciones universitarias,puntos porcentuales.Made for Freddy Vega, by HanademiFuentes: Chirikov, I. (2026). Study of AI exposure and more than 500,000 university grades, summarized by The Decoder.Unidadpuntos porcentuales16 %Cursos con muchas tareas13 %Todos los cursos
En más de medio millón de calificaciones, la proporción de sobresalientes subió 13 puntos tras ChatGPT. En cursos con muchas tareas, el aumento llegó a 16 puntos. Las mejores notas son visibles, pero este diseño no demuestra que el conocimiento subyacente mejorara.
AI detectors are too inaccurate to restoretrustOverall accuracy on 192 human, AI-generated and mixed academic texts in a 2026 evaluation.Made for Freddy Vega, by HanademiSources: Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academiccontexts. International Journal for Educational Integrity, 22, Article 4.Hanademi analysis: Assessment use reached eighty-eight percent, the two independently tested detectors imply error rates of thirty-one tothirty-nine percent, and the Brown scores differed by forty-seven point four points between the suspected take-home midterm and controlled final.Unitaccuracy score0.7Originality0.6Turnitin
A 2026 academic test gave Originality an overall accuracy of 0.69 and Turnitin 0.61. Both struggled with mixed human-AI writing. Institutions cannot rebuild assessment trust by asking an uncertain detector to make high-stakes judgments.
Los detectores de IA son demasiadoimprecisos para restaurar la confianzaPrecisión general en 192 textos académicos humanos, generados por IA y mixtos en una evaluación de2026.Made for Freddy Vega, by HanademiFuentes: Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors inacademic contexts. International Journal for Educational Integrity, 22, Article 4.Análisis de Hanademi: El uso de IA en evaluaciones llegó a ochenta y ocho por ciento, los dos detectores probados de forma independiente implican tasasde error de treinta y uno a treinta y nueve por ciento, y las notas de Brown difirieron en cuarenta y siete coma cuatro puntos entre el parcial…Unidadíndice de precisión0,7Originality0,6Turnitin
Una prueba académica de 2026 dio a Originality una precisión general de 0,69 y a Turnitin de 0,61. Ambos tuvieron dificultades con textos mixtos de humanos e IA. Las instituciones no pueden reconstruir la confianza pidiendo a un detector incierto que tome decisiones de alto riesgo.
Assessment must measure thestudent, not the submission.This synthesis combines adoption, detector performance and the Brown assessmentreversal into an assessment-design conclusion.Sources: Freeman, J. (2025). Student Generative AI Survey 2025. Higher Education Policy Institute and Kortext.; Hadra, M., Cambridge, K., & Mesbah, M.(2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, Article 4.Hanademi analysis: Assessment use reached eighty-eight percent, the two independently tested detectors imply error rates of thirty-one to thirty-nine percent,and the Brown scores differed by forty-seven point four points between the suspected take-home midterm and controlled final.
Mass adoption makes take-home output easier to produce and harder to authenticate. Detection remains imperfect, and a polished submission may no longer reveal who can reproduce the reasoning. The durable response is to redesign what students must demonstrate.
La evaluación debe medir al alumno,no la entrega.Esta síntesis combina adopción, desempeño de detectores y la reversión de Brown enuna conclusión sobre diseño de evaluaciones.Fuentes: Freeman, J. (2025). Student Generative AI Survey 2025. Higher Education Policy Institute and Kortext.; Hadra, M., Cambridge, K., & Mesbah, M.(2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, Article 4.Análisis de Hanademi: El uso de IA en evaluaciones llegó a ochenta y ocho por ciento, los dos detectores probados de forma independiente implican tasas de error de treinta y uno atreinta y nueve por ciento, y las notas de Brown difirieron en cuarenta y siete coma cuatro puntos entre el parcial sospechoso para casa y el final controlado.
La adopción masiva facilita producir tareas para casa y dificulta autentificarlas. La detección sigue siendo imperfecta y una entrega pulida puede dejar de mostrar quién puede reproducir el razonamiento. La respuesta duradera es rediseñar lo que los alumnos deben demostrar.
Works now and learned for later aredifferent outcomesShort-term AI-feedback effects and later proctored retention effects from separate large-scale mathstudies, percent change.Made for Freddy Vega, by HanademiSources: Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). Faster completion, less learning. arXiv.; A large scale randomizedcontrol trial showing LLM generated feedback helps low-knowledge middle school math students with short-term learning. (2026). ACM Digital Library.Hanademi analysis: One large experiment found wrong-answer correction up sixteen percent and next-problem success up seven percent, while aseparate dataset of three point two million interactions found a twenty-five-percent decline in the odds of later proctored retention success.Unitpercent changeWrong-answer correction16%Next-problem success7%Later retention odds-25%
One trial found that AI feedback made wrong answers 16% more likely to be corrected and improved the next problem by 7%. A separate dataset linked AI-susceptible problems to a 25% decline in later proctored retention odds. A tool can help on the next step without guaranteeing durable learning.
Funciona ahora y aprendido para despuésson resultados distintosEfectos a corto plazo de la retroalimentación con IA y efectos posteriores en retención supervisada deestudios matemáticos distintos, cambio porcentual.Made for Freddy Vega, by HanademiFuentes: Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). Faster completion, less learning. arXiv.; A large scale randomizedcontrol trial showing LLM generated feedback helps low-knowledge middle school math students with short-term learning. (2026). ACM Digital Library.Análisis de Hanademi: Un gran experimento encontró un aumento de dieciséis por ciento en la corrección de respuestas erróneas y de siete por ciento enel éxito del problema siguiente, mientras otro conjunto de tres coma dos millones de interacciones halló una caída de veinticinco por ciento en las…Unidadcambio porcentualCorrección de respuesta16 %Éxito en el siguiente problema7 %Retención posterior-25 %
Un ensayo halló que la retroalimentación con IA aumentó 16% la corrección de respuestas erróneas y mejoró 7% el siguiente problema. Otro conjunto de datos vinculó los problemas vulnerables a la IA con una caída del 25% en la probabilidad de retención supervisada posterior. Una herramienta puede ayudar en el siguiente paso sin garantizar aprendizaje duradero.
Students spent less time on AI-susceptiblemathChange in time spent on AI-susceptible math problems after ChatGPT, by education level.Made for Freddy Vega, by HanademiSources: Rismanchian, S., et al. (2026). Faster completion, less learning: Generative AI reduced study time on math problems and theknowledge they build. arXiv.Unitpercent changeHigh school-31.3%College-26.9%Middle school-9%
After ChatGPT, time spent on susceptible math problems fell at every education level studied. The decline reached 31.3% in high school and 26.9% in college. AI can act like an escalator: arriving faster does not show whether the rider’s legs became stronger.
Los estudiantes dedicaron menos tiempo amatemáticas vulnerables a la IACambio en el tiempo dedicado a problemas matemáticos vulnerables a la IA tras ChatGPT, por niveleducativo.Made for Freddy Vega, by HanademiFuentes: Rismanchian, S., et al. (2026). Faster completion, less learning: Generative AI reduced study time on math problems andthe knowledge they build. arXiv.Unidadcambio porcentualSecundaria-31,3 %Universidad-26,9 %Enseñanza media-9 %
Tras ChatGPT, el tiempo dedicado a problemas matemáticos vulnerables cayó en todos los niveles estudiados. La reducción alcanzó 31,3% en secundaria y 26,9% en universidad. La IA puede actuar como una escalera mecánica: llegar antes no demuestra que las piernas del usuario se fortalecieran.
Writing with complete AI relianceweakened comprehension mostChange in comprehension accuracy across three cycles of AI-assisted reading and writing tasks.Made for Freddy Vega, by HanademiSources: Duke University Innovation Co-Lab. (2023, September 27). Grant update: Impact of AI tools on reader comprehension.Unitpercent change-12%AI-assisted reading-25.1%AI writing
The strongest decline appeared when AI took over the writing process. Comprehension accuracy fell 25.1%, compared with 12% during AI-assisted reading. The more completely the tool substituted for producing the argument, the weaker the measured understanding.
Escribir con dependencia total de la IAdebilitó más la comprensiónCambio en la precisión de comprensión durante tres ciclos de lectura y escritura asistidas por IA.Made for Freddy Vega, by HanademiFuentes: Duke University Innovation Co-Lab. (2023, September 27). Grant update: Impact of AI tools on reader comprehension.Unidadcambio porcentual-12 %Lectura asistida por IA-25,1 %Escritura con IA
La mayor caída apareció cuando la IA asumió el proceso de escritura. La precisión de comprensión bajó 25,1%, frente al 12% durante la lectura asistida por IA. Cuanto más completamente sustituyó la herramienta la producción del argumento, más débil fue la comprensión medida.
The brain rot headline rests on the smalleststudyParticipant counts in the MIT EEG sessions versus large behavioral learning studies.Made for Freddy Vega, by HanademiSources: Kosmyna, N., et al. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writingtask. MIT Media Lab.; Strömberg, D., Lei, V., & Wu, Y. (2026). The generative AI learning penalty. CEPR Discussion Paper 21577.Hanademi analysis: The neural experiment’s fifty-four initial participants compare with twenty thousand seven hundred six students in the feedback trial andtwenty-six thousand eight hundred eleven in the homework-and-exam study, making those samples about three hundred eighty-three and four hundred…UnitparticipantsEEG switch session18EEG main sessions54AI feedback trial20,706Homework and exams26,811
The widely shared EEG experiment began with only 54 participants, and 18 completed its tool-switching session. The large behavioral studies here involve more than 20,000 students. Lower neural connectivity during one essay task is worth studying, but it is not proof of damaged brains or lower intelligence.
El titular sobre podredumbre cerebraldescansa en el estudio más pequeñoNúmero de participantes en las sesiones EEG del MIT frente a grandes estudios conductuales deaprendizaje.Made for Freddy Vega, by HanademiFuentes: Kosmyna, N., et al. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writingtask. MIT Media Lab.; Strömberg, D., Lei, V., & Wu, Y. (2026). The generative AI learning penalty. CEPR Discussion Paper 21577.Análisis de Hanademi: Los cincuenta y cuatro participantes iniciales del experimento neuronal se comparan con veinte mil setecientos seis estudiantes del ensayode retroalimentación y veintiséis mil ochocientos once del estudio de tareas y exámenes, lo que hace esas muestras unas trescientas ochenta y tres y…UnidadparticipantesSesión EEG de cambio18Sesiones EEG principales54Ensayo de retroalimentación20.706Tareas y exámenes26.811
El difundido experimento EEG comenzó con solo 54 participantes y 18 completaron la sesión de cambio de herramienta. Los grandes estudios conductuales aquí superan los 20.000 estudiantes. Una menor conectividad durante una tarea de escritura merece estudio, pero no demuestra daño cerebral ni menor inteligencia.
The strong K-12 evidence base is tinyStanford review of K-12 AI research, from papers identified to studies with strong causal evidence.Made for Freddy Vega, by HanademiSources: Fesler, L., Martinez Claeys, J. P., Agnew, C., & Loeb, S. (2026). The evidence base on AI in K-12: A 2026 review. Stanford SCALEInitiative.UnitpapersPapers found · 800Strong causal evidence · 20Strong U.S. evidence · 0
The volume of AI education research looks reassuring until quality filters are applied. Stanford found more than 800 papers, but only 20 offered strong cause-and-effect evidence. None met that standard for U.S. students, so confident universal claims are premature.
La base sólida de pruebas escolares esdiminutaRevisión de Stanford sobre investigación de IA escolar, desde trabajos identificados hasta estudios conpruebas causales sólidas.Made for Freddy Vega, by HanademiFuentes: Fesler, L., Martinez Claeys, J. P., Agnew, C., & Loeb, S. (2026). The evidence base on AI in K-12: A 2026 review. Stanford SCALEInitiative.UnidadtrabajosTrabajos encontrados · 800Pruebas causales sólidas · 20Pruebas sólidas de EE. UU. · 0
El volumen de investigación educativa sobre IA parece tranquilizador hasta aplicar filtros de calidad. Stanford encontró más de 800 trabajos, pero solo 20 ofrecían pruebas causales sólidas. Ninguno alcanzó ese nivel para estudiantes de Estados Unidos, por lo que las afirmaciones universales siguen siendo prematuras.
The boundary is substitutionversus scaffolding.This synthesis distinguishes AI that completes cognitive work from AI that structurespractice, feedback and reflection.Sources: ChatGPT’s impact on student learning outcomes: A meta-analysis of 35 experimental studies. (2026). Humanities and Social Sciences Communications.;Hong, H., Vate-U-Lan, P., & Viriyavejakul, C. (2025). Enhancing critical thinking. Proceedings of the International Conference on AI-enabled Education.Hanademi analysis: Complete reliance for writing reduced comprehension by twenty-five point one percent, while a structured method reduced unnecessary effort bythirty-two percent and increased learning effort by twenty-eight percent, consistent with a thirty-five-experiment synthesis finding moderate but design-dependent gains.
The same technology can remove thinking or redirect effort toward it. Unstructured answer generation encourages substitution. Good tutoring withholds the lift, gives feedback and leaves the student responsible for the reasoning.
La frontera está entre sustitución y apoyo.Esta síntesis distingue la IA que completa el trabajo cognitivo de la que estructurapráctica, retroalimentación y reflexión.Fuentes: ChatGPT’s impact on student learning outcomes: A meta-analysis of 35 experimental studies. (2026). Humanities and Social Sciences Communications.;Hong, H., Vate-U-Lan, P., & Viriyavejakul, C. (2025). Enhancing critical thinking. Proceedings of the International Conference on AI-enabled Education.Análisis de Hanademi: La dependencia completa para escribir redujo la comprensión en veinticinco coma uno por ciento, mientras un método estructurado redujo el esfuerzoinnecesario en treinta y dos por ciento y aumentó el esfuerzo de aprendizaje en veintiocho por ciento, en consonancia con una síntesis de treinta y cinco experimentos que…
La misma tecnología puede eliminar el pensamiento o redirigir el esfuerzo hacia él. La generación de respuestas sin estructura favorece la sustitución. Una buena tutoría no levanta el peso, ofrece retroalimentación y mantiene al alumno responsable del razonamiento.
Across 35 experiments, ChatGPTmoderately improved learning, dependingon teaching design.The synthesis covered experiments published from 2022 through 2024 and does notimply that every implementation helps.Sources: ChatGPT’s impact on student learning outcomes: A meta-analysis of 35 experimental studies. (2026). Humanities and Social SciencesCommunications.
The evidence is not uniformly negative. A 2026 synthesis of 35 experiments and 4,193 participants found a moderate overall learning benefit. The result changed with subject, duration and instructional mode, making pedagogy the active ingredient.
En 35 experimentos, ChatGPT mejorómoderadamente el aprendizaje, segúnel diseño docente.La síntesis cubrió experimentos publicados entre 2022 y 2024 y no implica que todaimplementación ayude.Fuentes: ChatGPT’s impact on student learning outcomes: A meta-analysis of 35 experimental studies. (2026). Humanities and Social SciencesCommunications.
Las pruebas no son uniformemente negativas. Una síntesis de 2026 de 35 experimentos y 4.193 participantes halló un beneficio moderado general. El resultado cambió según la materia, la duración y el modo de enseñanza, lo que convierte la pedagogía en el ingrediente activo.
Six supported weeks delivered years oflearningEstimated learning gains from six weeks of teacher-supported AI tutoring for first-year secondary studentsin Nigeria.Made for Freddy Vega, by HanademiSources: De Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W., & Dikoru, E. J. (2025). Fromchalkboards to chatbots. World Bank.Hanademi analysis: A randomized laptop evaluation covering five hundred thirty-one schools found no significant academic gains over tenyears, whereas six weeks of teacher-supported AI tutoring produced gains valued at one and a half to two years of ordinary schooling.Unitequivalent years of schooling2High estimate1.5Low estimate
This is the hopeful counterexample. Six weeks of teacher-supported tutoring produced gains estimated at 1.5 to 2 years of ordinary schooling. The device was not the intervention by itself; the structure, teacher and tutoring rules were part of the treatment.
Seis semanas con apoyo produjeron años deaprendizajeAvances estimados tras seis semanas de tutoría con IA y apoyo docente para alumnos del primer año desecundaria en Nigeria.Made for Freddy Vega, by HanademiFuentes: De Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W., & Dikoru, E. J. (2025). Fromchalkboards to chatbots. World Bank.Análisis de Hanademi: Una evaluación aleatoria de portátiles que cubrió quinientas treinta y una escuelas no encontró avancesacadémicos significativos durante diez años, mientras seis semanas de tutoría con IA apoyada por docentes produjeron mejoras…Unidadaños equivalentes de escolarización2Estimación alta1,5Estimación baja
Este es el contraejemplo esperanzador. Seis semanas de tutoría con apoyo docente produjeron avances estimados entre 1,5 y 2 años de escolarización habitual. El dispositivo no fue la intervención por sí solo; la estructura, el docente y las reglas tutoriales formaron parte del tratamiento.
AI helped most where human support hadmore room to improveMath mastery gains in a randomized trial involving 900 tutors and 1,800 underserved K-12 students,percentage points.Made for Freddy Vega, by HanademiSources: Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024). Tutor CoPilot: A human-AI approach forscaling real-time expertise. Stanford University.Hanademi analysis: Students with lower-rated tutors gained nine mastery points versus four points overall, so their gain was two andone-quarter times as large, while teacher-supported tutoring elsewhere delivered gains valued at one and a half to two ordinary school years.Unitpercentage pointsAll students4%Lower-rated tutors9%2.2x
Tutor CoPilot raised mastery by 4 points overall. For students working with lower-rated tutors, the gain reached 9 points. AI’s strongest promise here is equalization, scaling better guidance through a human relationship rather than replacing it.
La IA ayudó más donde el apoyo humanotenía más margen de mejoraAvances en dominio matemático en un ensayo aleatorizado con 900 tutores y 1.800 estudiantes escolaresdesfavorecidos, puntos porcentuales.Made for Freddy Vega, by HanademiFuentes: Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024). Tutor CoPilot: A human-AI approach forscaling real-time expertise. Stanford University.Análisis de Hanademi: Los estudiantes con tutores de menor calificación ganaron nueve puntos de dominio frente a cuatro puntos en general, por lo quesu mejora fue dos coma veinticinco veces mayor, mientras la tutoría apoyada por docentes en otro contexto produjo avances valorados entre uno…Unidadpuntos porcentualesTodos los alumnos4 %Tutores con menor valoración9 %2,2x
Tutor CoPilot elevó el dominio 4 puntos en general. Entre alumnos con tutores de menor valoración, el avance llegó a 9 puntos. La mayor promesa de la IA aquí es la igualación, ampliando una mejor orientación mediante una relación humana en vez de reemplazarla.
AI tutoring needs subject-specific qualitycontrolsChatGPT help failing quality checks before and after consistency filtering across mathematics areas.Made for Freddy Vega, by HanademiSources: Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored help onmathematics skills. PLOS ONE.Hanademi analysis: Seventy-three percent finding AI clarifying minus forty-three percent finding it reliable gives a thirty-point gap, whilefiltering cut an initial thirty-two-percent failure rate to near zero in algebra but only to thirteen percent in statistics.Unitfailure rateQUALITY-CHECK STAGEFAILURE RATEInitial help32%Algebra after filteringNear 0%Statistics after filtering13%
Raw ChatGPT help failed quality checks on 32% of problems. Consistency filtering reduced failures close to zero in algebra, but only to 13% in statistics. Fluency can sound equally convincing across subjects even when reliability differs.
La tutoría con IA necesita controlesespecíficos por materiaAyuda de ChatGPT que falló controles de calidad antes y después del filtrado de consistencia en áreasmatemáticas.Made for Freddy Vega, by HanademiFuentes: Pardos, Z. A., & Bhandari, S. (2024). ChatGPT-generated help produces learning gains equivalent to human tutor-authored helpon mathematics skills. PLOS ONE.Análisis de Hanademi: Setenta y tres por ciento que consideró aclaratoria a la IA menos cuarenta y tres por ciento que la consideró fiable produce una brechade treinta puntos, mientras el filtrado redujo una tasa inicial de fallos de treinta y dos por ciento a casi cero en álgebra pero solo a trece por ciento en…Unidadtasa de fallosETAPA DEL CONTROLTASA DE FALLOSAyuda inicial32%Álgebra tras filtradoCasi 0%Estadística tras filtrado13%
La ayuda inicial de ChatGPT falló controles de calidad en el 32% de los problemas. El filtrado redujo los fallos casi a cero en álgebra, pero solo al 13% en estadística. La fluidez puede sonar igual de convincente entre materias aunque la fiabilidad difiera.
Good offloading removes friction withoutremoving thoughtChange in unnecessary mental effort and learning-oriented effort under a structured AI-writing method.Made for Freddy Vega, by HanademiSources: Hong, H., Vate-U-Lan, P., & Viriyavejakul, C. (2025). Enhancing critical thinking: Interactive cognitive offload instruction withgenerative AI in English essay writing. Proceedings of the International Conference on AI-enabled Education.Unitpercent changeUnnecessary effort-32%Learning effort28%
Cognitive offloading is not automatically harmful. This structured method delegated mechanical writing work while requiring reflection, evidence checking and student control. It reduced unnecessary effort by 32% and redirected effort toward learning by 28%.
Una buena delegación elimina fricción sineliminar pensamientoCambio en el esfuerzo mental innecesario y el esfuerzo orientado al aprendizaje con un método estructuradode escritura con IA.Made for Freddy Vega, by HanademiFuentes: Hong, H., Vate-U-Lan, P., & Viriyavejakul, C. (2025). Enhancing critical thinking: Interactive cognitive offload instructionwith generative AI in English essay writing. Proceedings of the International Conference on AI-enabled Education.Unidadcambio porcentualEsfuerzo innecesario-32 %Esfuerzo de aprendizaje28 %
La delegación cognitiva no es automáticamente dañina. Este método estructurado delegó trabajo mecánico de escritura, pero exigió reflexión, verificación de pruebas y control estudiantil. Redujo 32% el esfuerzo innecesario y redirigió 28% más esfuerzo hacia el aprendizaje.
Ten years of laptops brought no significantgains in rural Peru.The long-run laptop result is a caution against treating any educational technology as aself-executing intervention.Sources: Cueto, S., Beuermann, D., Cristia, J., Malamud, O., & Ramos Pardo, F. J. (2025). Laptops in the long run: Evidence from the One Laptop perChild program in rural Peru. NBER Working Paper 34495.Hanademi analysis: A randomized laptop evaluation covering five hundred thirty-one schools found no significant academic gains over ten years, whereas sixweeks of teacher-supported AI tutoring produced gains valued at one and a half to two years of ordinary schooling.
A randomized evaluation followed 531 rural Peruvian schools for ten years. Laptops improved computer skills but produced no significant academic gains and some evidence of worse grade progression. Technology access matters, but classroom use and teacher preparation determine whether access becomes learning.
Diez años de portátiles no produjeronningún avance significativo en el Perú rural.El resultado de largo plazo con portátiles advierte contra tratar cualquier tecnologíaeducativa como una intervención automática.Fuentes: Cueto, S., Beuermann, D., Cristia, J., Malamud, O., & Ramos Pardo, F. J. (2025). Laptops in the long run: Evidence from the One Laptop perChild program in rural Peru. NBER Working Paper 34495.Análisis de Hanademi: Una evaluación aleatoria de portátiles que cubrió quinientas treinta y una escuelas no encontró avances académicos significativos durante diezaños, mientras seis semanas de tutoría con IA apoyada por docentes produjeron mejoras valoradas entre uno coma cinco y dos años de escolaridad ordinaria.
Una evaluación aleatorizada siguió durante diez años a 531 escuelas rurales de Perú. Los portátiles mejoraron habilidades informáticas, pero no produjeron avances académicos significativos y hubo indicios de peor progresión escolar. El acceso importa, pero el uso en clase y la preparación docente determinan si se convierte en aprendizaje.
The PISA collapse began before ChatGPTChange in OECD average PISA performance from 2018 to 2022, before widespread public generative AIadoption.Made for Freddy Vega, by HanademiSources: OECD. (2023). PISA 2022 results, Volume I: The state of learning and equity in education. OECD Publishing.Hanademi analysis: From twenty eighteen to twenty twenty-two, United States internet use rose from about eighty-eight point five to ninety-two pointseven percent, a gain of four point two points, while OECD mathematics fell fifteen PISA points and reading fell ten before widespread ChatGPT use.UnitPISA points-10Reading-15Mathematics
The education crisis did not start with generative AI. Between 2018 and 2022, OECD mathematics fell 15 points and reading fell 10. PISA 2022 was largely administered before public ChatGPT adoption, so blaming the original collapse on chatbots gets the chronology wrong.
El desplome de PISA comenzó antes deChatGPTCambio en el rendimiento promedio de PISA en la OCDE entre 2018 y 2022, antes de la adopción públicageneralizada de IA generativa.Made for Freddy Vega, by HanademiFuentes: OECD. (2023). PISA 2022 results, Volume I: The state of learning and equity in education. OECD Publishing.Análisis de Hanademi: De dos mil dieciocho a dos mil veintidós, el uso de internet en Estados Unidos subió de aproximadamente ochenta y ocho coma cinco anoventa y dos coma siete por ciento, un aumento de cuatro coma dos puntos, mientras las matemáticas de la OCDE cayeron quince puntos PISA y la lectura diez…Unidadpuntos PISA-10Lectura-15Matemáticas
La crisis educativa no comenzó con la IA generativa. Entre 2018 y 2022, matemáticas cayó 15 puntos en la OCDE y lectura bajó 10. PISA 2022 se aplicó en gran medida antes de la adopción pública de ChatGPT, por lo que culpar a los chatbots del desplome original contradice la cronología.
One product cycle rewrote codingexpectationsFrontier-model success on SWE-bench real software issues, 2023 and 2024.Made for Freddy Vega, by HanademiSources: Stanford HAI. (2025). Technical performance. 2025 AI Index Report.Hanademi analysis: The coding solve rate rose from four point four to seventy-one point seven percent, which is sixteen point three times ashigh, while a seven-month doubling compounds to about three point three times over the twelve months between annual labels.Unitpercent solved20234.4%202471.7%
In 2023, frontier models solved 4.4% of SWE-bench problems. In 2024, they solved 71.7%. The benchmark covers real software issues, not every workplace task, but the speed means curriculum assumptions can expire inside one academic year.
Un ciclo de producto reescribió lasexpectativas de programaciónÉxito de modelos avanzados en problemas reales de software de SWE-bench, 2023 y 2024.Made for Freddy Vega, by HanademiFuentes: Stanford HAI. (2025). Technical performance. 2025 AI Index Report.Análisis de Hanademi: La tasa de resolución de problemas de programación subió de cuatro coma cuatro a setenta y uno coma siete por ciento, lo queequivale a dieciséis coma tres veces, mientras una duplicación cada siete meses se acumula hasta unas tres coma tres veces durante los doce meses…Unidadporcentaje resuelto20234,4 %202471,7 %
En 2023, los modelos avanzados resolvían el 4,4% de los problemas de SWE-bench. En 2024, resolvían el 71,7%. El indicador cubre problemas reales de software, no todas las tareas laborales, pero la velocidad implica que los supuestos curriculares pueden caducar dentro de un año académico.
AI’s reliable software-taskhorizon has doubled about everyseven months since 2019.The measure is limited to software tasks and human-expert completion times;extrapolation to other domains remains uncertain.Sources: METR. (2025, March 19). Measuring AI ability to complete long tasks.Hanademi analysis: The coding solve rate rose from four point four to seventy-one point seven percent, which is sixteen point three times as high, while aseven-month doubling compounds to about three point three times over the twelve months between annual labels.
Capability is not only improving on short benchmarks. Since 2019, the duration of software tasks frontier AI can complete with 50% reliability has doubled about every seven months. Education must prepare students for moving boundaries, not freeze training around today’s division of labor.
El horizonte fiable de tareas de software dela IA se duplica aproximadamente cadasiete meses desde 2019.La medida se limita a tareas de software y tiempos de expertos humanos; extrapolarla aotros ámbitos sigue siendo incierto.Fuentes: METR. (2025, March 19). Measuring AI ability to complete long tasks.Análisis de Hanademi: La tasa de resolución de problemas de programación subió de cuatro coma cuatro a setenta y uno coma siete por ciento, lo que equivale a dieciséiscoma tres veces, mientras una duplicación cada siete meses se acumula hasta unas tres coma tres veces durante los doce meses entre las etiquetas anuales.
La capacidad no mejora solo en pruebas breves. Desde 2019, la duración de las tareas de software que la IA avanzada completa con un 50% de fiabilidad se duplica aproximadamente cada siete meses. La educación debe preparar a los alumnos para fronteras móviles y no congelar la formación en la división actual del trabajo.
A net jobs gain still hides massive churnWorld Economic Forum employer forecast for jobs created, displaced and net change by 2030, millions; notan AI-only forecast.Made for Freddy Vega, by HanademiSources: World Economic Forum. (2025, January 7). Future of Jobs Report 2025: 78 million new job opportunities by 2030 but urgentupskilling needed.Hanademi analysis: Adding one hundred seventy million projected jobs created to ninety-two million displaced gives two hundred sixty-two million in gross churn,which is three point four times the seventy-eight-million net gain, as Gen Z anger rose nine points while excitement fell fourteen and hopefulness fell nine.Unitmillion jobs050100150170Created-92Displaced78Net change
The headline is positive: 170 million jobs created, 92 million displaced and a net gain of 78 million by 2030. The lived transition is much larger than the net. Creation plus displacement puts 262 million jobs inside the churn, even before asking who can move between them.
Una ganancia neta de empleo aún oculta unaenorme rotaciónPrevisión empresarial del Foro Económico Mundial sobre empleos creados, desplazados y cambio neto para2030, millones; no es una previsión exclusiva de la IA.Made for Freddy Vega, by HanademiFuentes: World Economic Forum. (2025, January 7). Future of Jobs Report 2025: 78 million new job opportunities by 2030 buturgent upskilling needed.Análisis de Hanademi: Sumar ciento setenta millones de empleos proyectados como creados y noventa y dos millones desplazados da doscientossesenta y dos millones de rotación bruta, equivalente a tres coma cuatro veces la ganancia neta de setenta y ocho millones, mientras la ira de la…Unidadmillones de empleos050100150170Creados-92Desplazados78Cambio neto
El titular es positivo: 170 millones de empleos creados, 92 millones desplazados y una ganancia neta de 78 millones para 2030. La transición vivida es mucho mayor que el saldo. La creación y el desplazamiento sitúan 262 millones de empleos dentro de la rotación, antes incluso de preguntar quién podrá pasar de unos a otros.
Gen Z is turning from excitement towardangerChange in U.S. Gen Z sentiment toward AI from 2025 to 2026, percentage points; Gallup survey of ages14-29.Made for Freddy Vega, by HanademiSources: Gallup. (2026, April 9). Gen Z's AI adoption steady, but skepticism climbs.; Walsh, M. (2026, April 17). Frustration,skepticism: Survey reveals shifting Gen Z attitudes toward AI. Education Week.Hanademi analysis: Adding one hundred seventy million projected jobs created to ninety-two million displaced gives two hundred sixty-two million in grosschurn, which is three point four times the seventy-eight-million net gain, as Gen Z anger rose nine points while excitement fell fourteen and hopefulness…Unitpercentage pointsAnger9%Excitement-14%Hopefulness-9%
Young people are not simply refusing AI. Their emotional response is shifting while adoption continues. Anger rose 9 points, excitement fell 14 and hopefulness fell 9, a pattern consistent with uncertainty about agency and future work.
La generación Z pasa del entusiasmo alenojoCambio en el sentimiento de la generación Z estadounidense ante la IA de 2025 a 2026, puntosporcentuales; encuesta de Gallup a personas de 14 a 29 años.Made for Freddy Vega, by HanademiFuentes: Gallup. (2026, April 9). Gen Z's AI adoption steady, but skepticism climbs.; Walsh, M. (2026, April 17). Frustration,skepticism: Survey reveals shifting Gen Z attitudes toward AI. Education Week.Análisis de Hanademi: Sumar ciento setenta millones de empleos proyectados como creados y noventa y dos millones desplazados da doscientossesenta y dos millones de rotación bruta, equivalente a tres coma cuatro veces la ganancia neta de setenta y ocho millones, mientras la ira de la…Unidadpuntos porcentualesEnojo9 %Entusiasmo-14 %Esperanza-9 %
Los jóvenes no están simplemente rechazando la IA. Su respuesta emocional cambia mientras la adopción continúa. El enojo subió 9 puntos, el entusiasmo cayó 14 y la esperanza bajó 9, un patrón compatible con incertidumbre sobre la autonomía y el trabajo futuro.
Young people also use AI to copeYouth reporting AI use for mental health support or personal conversation and advice in New South Wales,2026.Made for Freddy Vega, by HanademiSources: New South Wales Government. (2026, April 13). From AI to anxiety: New poll reveals the state of young people in2026.Hanademi analysis: Twenty-nine percent of youth used AI for mental-health support and twenty-seven percent for conversation orpersonal advice, only two points apart, even as anger rose nine points and both excitement and hopefulness declined.Unitpercent of respondentsMental health supportConversation or advice1.1×29%27%
In a 2026 youth poll, 29% used AI to support their mental health and 27% used it for conversation or personal advice. That creates an unusual feedback loop. The technology young people fear may also become the place where they process that fear.
Los jóvenes también usan la IA paraafrontar sus problemasJóvenes que declararon usar IA para apoyo de salud mental o conversación y consejos personales en NuevaGales del Sur, 2026.Made for Freddy Vega, by HanademiFuentes: New South Wales Government. (2026, April 13). From AI to anxiety: New poll reveals the state of young peoplein 2026.Análisis de Hanademi: Veintinueve por ciento de los jóvenes usó IA como apoyo para la salud mental y veintisiete por ciento para conversar orecibir consejos personales, una diferencia de solo dos puntos, mientras la ira subió nueve puntos y tanto el entusiasmo como la esperanza…Unidadporcentaje de encuestadosApoyo de salud mentalConversación o consejos1,1×29 %27 %
En una encuesta juvenil de 2026, el 29% usó IA para apoyar su salud mental y el 27% para conversar o pedir consejos personales. Esto crea un bucle inusual. La tecnología que los jóvenes temen también puede convertirse en el lugar donde procesan ese temor.
Internet access soared while youthunemployment ended near its starting pointU.S. internet use and modeled youth unemployment, 2000-2024; historical context, not a forecast or causaltest.Made for Freddy Vega, by HanademiSources: World Bank. (n.d.). Individuals using the Internet (% of population) [IT.NET.USER.ZS]. World Bank Open Data.; World Bank. (n.d.).Unemployment, youth total (% of total labor force ages 15-24) (modeled ILO estimate) [SL.UEM.1524.ZS]. World Bank Open Data.Hanademi analysis: Between two thousand and twenty twenty-four, United States internet use rose from forty-three point one to ninety-four point sevenpercent, a gain of fifty-one point six points, while youth unemployment moved from nine point three to eight point nine percent, down about four tenths of a point.UnitpercentInternet use43.1%94.7%20002024Youth unemployment9.3%8.9%20002024
Internet use rose from 43.1% in 2000 to 94.7% in 2024. Youth unemployment moved through recessions and the pandemic but ended at 8.92%, close to and slightly below its 2000 level. This is not proof that AI disruption will be harmless, only a reminder that technological diffusion does not mechanically determine the employment endpoint.
El acceso a internet se disparó mientras eldesempleo juvenil terminó cerca del inicioUso de internet y desempleo juvenil modelado en Estados Unidos, 2000-2024; contexto histórico, no unaprevisión ni una prueba causal.Made for Freddy Vega, by HanademiFuentes: World Bank. (n.d.). Individuals using the Internet (% of population) [IT.NET.USER.ZS]. World Bank Open Data.; World Bank. (n.d.).Unemployment, youth total (% of total labor force ages 15-24) (modeled ILO estimate) [SL.UEM.1524.ZS]. World Bank Open Data.Análisis de Hanademi: Entre dos mil y dos mil veinticuatro, el uso de internet en Estados Unidos subió de cuarenta y tres coma uno a noventa y cuatro coma siete porciento, un aumento de cincuenta y uno coma seis puntos, mientras el desempleo juvenil pasó de nueve coma tres a ocho coma nueve por ciento, una disminución de…UnidadporcentajeUso de internet43,1 %94,7 %20002024Desempleo juvenil9,3 %8,9 %20002024
El uso de internet subió del 43,1% en 2000 al 94,7% en 2024. El desempleo juvenil atravesó recesiones y la pandemia, pero terminó en 8,92%, cerca y ligeramente por debajo de su nivel de 2000. Esto no demuestra que la disrupción de la IA será inocua, solo recuerda que la difusión tecnológica no determina mecánicamente el resultado laboral.

The research behind this deck

AI can improve visible work while weakening independent performance, yet structured tutoring can reverse that tradeoff. The evidence points beyond bans or detection toward assessments, teaching methods and curricula designed around what students can still do alone.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

La IA puede mejorar el trabajo visible y debilitar el desempeño independiente, pero una tutoría estructurada puede invertir esa tensión. Las pruebas apuntan más allá de prohibiciones o detectores, hacia evaluaciones, métodos docentes y currículos centrados en lo que el alumno aún puede hacer solo.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English