Hanademi

When AI does the homeworkCuando la IA hace la tarea · V7

35 slides · 23 min · 2026-08-07 language
ENES
theme
LightDark
view
TalkTable
brand
HanademiPlatzi
When AI does the homeworkMade for Hugo L. Alvarez, by Hanademi
  1. The question is not whether AI improves an assignment, but whether the student can still perform alone.
  2. Unrestricted GPT users scored 17% below controls after the assistance disappeared.
  3. Universities should assess reasoning, revision, and verification instead of treating the final answer as sufficient evidence.
Cuando la IA hace la tareaMade for Hugo L. Alvarez, by Hanademi
  1. La pregunta no es si la IA mejora una tarea, sino si el estudiante todavía puede rendir por sí solo.
  2. Los usuarios de GPT sin restricciones puntuaron un 17% por debajo del control cuando desapareció la ayuda.
  3. Las universidades deben evaluar el razonamiento, la revisión y la verificación en vez de considerar suficiente la respuesta final.
Safeguarded AI delivered a 127% practicegainRelative mathematics practice improvement versus control in a Turkish high-school experiment.Made for Hugo L. Alvarez, by HanademiSources: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. NBERWorking Paper No. 33097.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.Unitrelative improvementUnrestricted GPT48%Safeguarded tutor127%2.6x
AI assistance is not one treatment. Unrestricted GPT raised practice performance 48%, while a tutor designed to require reasoning raised it 127%. The safeguarded tutor achieved that gain without the significant unaided penalty observed with unrestricted GPT.
La IA protegida logró una mejora prácticadel 127%Mejora relativa en la práctica de matemáticas frente al control en un experimento de secundaria enTurquía.Made for Hugo L. Alvarez, by HanademiFuentes: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. NBERWorking Paper No. 33097.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.Unidadmejora relativaGPT sin restricciones48 %Tutor protegido127 %2,6x
La ayuda de IA no es un tratamiento único. GPT sin restricciones elevó el rendimiento práctico un 48%, mientras un tutor diseñado para exigir razonamiento lo elevó un 127%. El tutor protegido logró esa mejora sin la penalización independiente significativa observada con GPT sin restricciones.
The terms behind the argumentMade for Hugo L. Alvarez, by HanademiGenerative AIArtificial intelligence that produces new text, images, audio, video, or code inresponse to instructions.Retrieval practiceAttempting to recall information without viewing the answer, therebystrengthening later access to that knowledge.MetacognitionPlanning, monitoring, and evaluating one's own thinking and learning.Competence frontierThe uneven boundary between tasks an AI system can perform reliably andsuperficially similar tasks where it fails.
4 technical terms carry much of the argument.
Los términos detrás del argumentoMade for Hugo L. Alvarez, by HanademiIA generativaInteligencia artificial que produce texto, imágenes, audio, video o código nuevos enrespuesta a instrucciones.Práctica de recuperaciónIntentar recordar información sin mirar la respuesta, fortaleciendoasí el acceso posterior a ese conocimiento.MetacogniciónPlanificar, supervisar y evaluar el propio pensamiento y aprendizaje.Frontera de competenciaEl límite irregular entre las tareas que un sistema de IA realiza confiabilidad y otras aparentemente similares en las que falla.
4 términos técnicos sostienen buena parte del argumento.
Unrestricted GPT turned a 48% practice gaininto a 17% unaided penaltyMade for Hugo L. Alvarez, by HanademiSources: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harmlearning. NBER Working Paper No. 33097.; Generative AI without guardrails can harm learning: Evidence from highschool mathematics.; Bastani, H., et al. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.UnitRelative performance change versus controlAssisted practiceUnrestricted GPT+48%Safeguarded tutor+127%After AI removalUnrestricted GPT-17%Safeguarded tutorNo significant disadvantage-2004080120Change versus control
Unrestricted GPT improved assisted work but left students below control when the tool disappeared. Safeguards delivered 2.6 times the practice gain without a significant unaided disadvantage.
GPT sin restricciones convirtió una mejora prácticadel 48% en una penalización autónoma del 17%Made for Hugo L. Alvarez, by HanademiFuentes: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning.NBER Working Paper No. 33097.; Generative AI without guardrails can harm learning: Evidence from high schoolmathematics.; Bastani, H., et al. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.UnidadCambio relativo del rendimiento frente al grupo de controlPráctica con asistenciaGPT libre+48%Tutor con salvaguardas+127%Después de retirar la IAGPT libre-17%Tutor con salvaguardasSin desventaja significativa-2004080120Cambio frente al grupo de control
GPT sin restricciones mejoró el trabajo asistido, pero dejó a los estudiantes por debajo del grupo de control al desaparecer la herramienta. Las salvaguardas lograron 2,6 veces la mejora práctica sin una desventaja autónoma significativa.
AI arrived early and evolved beforegraduationMonths from a typical August 2022 ninth-grade start to public ChatGPT, and from ChatGPT to GPT-4ocapabilities.Made for Hugo L. Alvarez, by HanademiSources: OpenAI. (2022). Introducing ChatGPT.; OpenAI. (2023). GPT-4 technical report.; OpenAI. (2024). Hello GPT-4o.UnitmonthsSchool start to ChatGPT3ChatGPT to GPT-4o186x
ChatGPT launched about 3 months after a typical August 2022 ninth-grade start. It then progressed from text-focused interaction to image, voice, and vision capabilities within roughly 18 months. The cohort did not have AI for every month, but it had access through nearly all four high-school years.
La IA llegó pronto y evolucionó antes de lagraduaciónMeses desde un inicio típico de noveno grado en agosto de 2022 hasta ChatGPT público, y desde ChatGPThasta las capacidades de GPT-4o.Made for Hugo L. Alvarez, by HanademiFuentes: OpenAI. (2022). Introducing ChatGPT.; OpenAI. (2023). GPT-4 technical report.; OpenAI. (2024). Hello GPT-4o.UnidadmesesInicio escolar a ChatGPT3ChatGPT a GPT-4o186x
ChatGPT se lanzó unos 3 meses después de un inicio típico de noveno grado en agosto de 2022. Luego pasó de una interacción centrada en texto a capacidades de imagen, voz y visión en aproximadamente 18 meses. La cohorte no tuvo IA durante todos los meses, pero sí durante casi los cuatro años de secundaria.
The 2026 entering class had ChatGPT for88% of high schoolDivide each cohort's years of public ChatGPT availability by a four-year high school span and multiply by100.Made for Hugo L. Alvarez, by HanademiSources: OpenAI. (2022). Introducing ChatGPT.; Introducing ChatGPT.UnitShare of four-year high school with public ChatGPT access02550752023 entrants2024 entrants2025 entrants2026 entrants12.587.5
Exposure rises by 25 percentage points with each entering class, from 12.5% for 2023 entrants to 87.5% for 2026 entrants.
La generación universitaria de 2026 tuvoChatGPT durante el 88% de la secundariaDividir los años de disponibilidad pública de ChatGPT para cada generación entre cuatro años de secundariay multiplicar por 100.Made for Hugo L. Alvarez, by HanademiFuentes: OpenAI. (2022). Introducing ChatGPT.; Introducing ChatGPT.UnidadProporción de cuatro años de secundaria con acceso público a ChatGPT0255075Ingreso de 2023Ingreso de 2024Ingreso de 2025Ingreso de 202612,587,5
La exposición aumenta 25 puntos porcentuales con cada generación de ingreso, desde el 12,5% para la de 2023 hasta el 87,5% para la de 2026.
Without AI, unrestricted-GPTstudents scored -17% relativeperformance versus controls.The comparison comes from the same Turkish high-school mathematics experiment asthe opener.Sources: Bastani, H., et al. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.; Generative AI without guardrails can harmlearning: Evidence from high school mathematics.
Practice performance looked better while unrestricted GPT was available. After the tool was removed, those students scored -17% relative performance versus controls. The final assisted answer had overstated what they could reproduce independently.
Sin IA, los estudiantes con GPT sinrestricciones obtuvieron un rendimientorelativo de -17% frente al control.La comparación procede del mismo experimento turco de matemáticas de secundaria queabre la presentación.Fuentes: Bastani, H., et al. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.; Generative AI without guardrails can harmlearning: Evidence from high school mathematics.
El rendimiento práctico parecía mejor mientras GPT sin restricciones estaba disponible. Al retirar la herramienta, esos estudiantes obtuvieron un rendimiento relativo de -17% frente al control. La respuesta asistida había exagerado lo que podían reproducir independientemente.
Teen ChatGPT use for schoolwork doubledin one yearShare of United States teenagers reporting ChatGPT use for schoolwork, 2023 and 2024.Made for Hugo L. Alvarez, by HanademiSources: Sidoti, O., & Gottfried, J. (2025). About a quarter of U.S. teens have used ChatGPT for schoolwork, double the share in2023. Pew Research Center.; pewresearch.org.Unitshare of teenagers26%202413%2023
Pew found that United States teen use of ChatGPT for schoolwork doubled from 13% in 2023 to 26% in 2024. By February 2026, just over half reported using chatbots for schoolwork. Another 12% had used them for emotional support, showing that chatbot use extends beyond assignments.
El uso adolescente de ChatGPT para tareasse duplicó en un añoProporción de adolescentes estadounidenses que reportó usar ChatGPT para tareas, 2023 y 2024.Made for Hugo L. Alvarez, by HanademiFuentes: Sidoti, O., & Gottfried, J. (2025). About a quarter of U.S. teens have used ChatGPT for schoolwork, double theshare in 2023. Pew Research Center.; pewresearch.org.Unidadproporción de adolescentes26 %202413 %2023
Pew encontró que el uso de ChatGPT para tareas entre adolescentes estadounidenses se duplicó del 13% en 2023 al 26% en 2024. En febrero de 2026, poco más de la mitad reportó usar chatbots para tareas. Otro 12% los había usado como apoyo emocional, lo que demuestra que su uso supera el ámbito académico.
The classroom is already digitalThree teenagers work together on laptop computers in a classroom.Sources: pewresearch.org.
The adoption numbers concern real teenagers already doing collaborative schoolwork on laptops.
El aula ya es digitalTres adolescentes trabajan juntos con computadoras portátiles en un aula.Fuentes: pewresearch.org.
Las cifras de adopción representan a adolescentes reales que ya realizan trabajo escolar colaborativo con computadoras portátiles.
AI use reached assessments, not just studysupportReported AI use among surveyed United Kingdom undergraduates in 2025.Made for Hugo L. Alvarez, by HanademiSources: Freeman, J. (2025). Student generative AI survey 2025. Higher Education Policy Institute.Unitshare of respondentsUsed AI · 92%Used in assessments · 88%Included generated text · 18%
HEPI reported that 92% of surveyed United Kingdom undergraduates used AI and 88% used it in assessments. Directly inserting generated text was much less common at 18%. Simple labels such as user and non-user hide major differences in how AI enters student work.
El uso de IA llegó a las evaluaciones, nosolo al estudioUso de IA declarado por universitarios encuestados en el Reino Unido en 2025.Made for Hugo L. Alvarez, by HanademiFuentes: Freeman, J. (2025). Student generative AI survey 2025. Higher Education Policy Institute.Unidadproporción de encuestadosUsó IA · 92 %Usó en evaluaciones · 88 %Incluyó texto generado · 18 %
HEPI informó que el 92% de los universitarios británicos encuestados usaba IA y el 88% la utilizaba en evaluaciones. Incluir directamente texto generado era mucho menos común, con un 18%. Etiquetas simples como usuario y no usuario ocultan diferencias importantes en la forma de usar IA.
A 2026 Georgia audit surveyedmore than 13,000 teachers.The audit establishes broad instructional presence, not a measured learning effect.Sources: audits.ga.gov.
Student adoption is only half the story. A June 2026 Georgia audit surveyed more than 13,000 public-school teachers and found generative AI was already part of instructional practice in many classrooms. AI exposure now comes through both personal tools and school activity.
Una auditoría de Georgia de 2026 encuestóa más de 13.000 docentes.La auditoría establece una presencia educativa amplia, no un efecto medido sobre elaprendizaje.Fuentes: audits.ga.gov.
La adopción estudiantil es solo la mitad de la historia. Una auditoría de Georgia de junio de 2026 encuestó a más de 13.000 docentes de escuelas públicas y encontró que la IA generativa ya formaba parte de la práctica educativa en muchas aulas. La exposición llega tanto por herramientas personales como por actividades escolares.
New Hampshire approved Khanmigo accessfor grades 5 through 12.The pilot establishes access across eligible grades, not a measured learning outcome.Sources: education.nh.gov.
AI assistance is becoming part of public education infrastructure. New Hampshire approved a statewide Khanmigo pilot for educators and students in grades 5 through 12. Universities should expect incoming students to have experienced very different tools and guardrails across schools.
New Hampshire aprobó acceso a Khanmigopara los grados 5 a 12.El piloto establece acceso en los grados elegibles, no un resultado medido de aprendizaje.Fuentes: education.nh.gov.
La ayuda de IA se está convirtiendo en parte de la infraestructura educativa pública. New Hampshire aprobó un piloto estatal de Khanmigo para docentes y estudiantes de los grados 5 a 12. Las universidades deben esperar experiencias muy distintas con herramientas y protecciones entre escuelas.
Across 936 tasks, higher AI confidencetracked lower critical effort.The study used self-reported workplace examples and establishes an association, not acausal learning effect.Sources: Lee, H. P., Sarkar, A., Tankelevitch, L., et al. (2025). The impact of generative AI on critical thinking. Proceedings of CHI 2025.
The critical-thinking risk is misplaced trust, not mere tool presence. Microsoft and Carnegie Mellon researchers analyzed 936 self-reported task examples from 319 knowledge workers. Higher confidence in AI was associated with less critical effort.
En 936 tareas, mayor confianza en IAacompañó menor esfuerzo crítico.El estudio utilizó ejemplos laborales declarados y establece una asociación, no un efectocausal sobre el aprendizaje.Fuentes: Lee, H. P., Sarkar, A., Tankelevitch, L., et al. (2025). The impact of generative AI on critical thinking. Proceedings of CHI 2025.
El riesgo para el pensamiento crítico es la confianza indebida, no la mera presencia de la herramienta. Investigadores de Microsoft y Carnegie Mellon analizaron 936 ejemplos de tareas declaradas por 319 trabajadores del conocimiento. Una mayor confianza en la IA se asoció con menor esfuerzo crítico.
AI made writing faster and betterChanges in writing time and independently rated quality in a 444-person preregistered experiment.Made for Hugo L. Alvarez, by HanademiSources: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.; Graham, S., &Perin, D. (2007). Writing next: Effective strategies to improve writing of adolescents in middle and high schools. Carnegie Corporation of New York.Sources use different measurement bases; read the comparison directionally, not as one exact scale.Unitpercent changeTime reduction40%Quality increase18%
ChatGPT reduced professional-writing time by 40% and raised independently rated quality by 18%. Yet pre-AI evidence found a 0.82 standardized writing effect associated with explicit strategy instruction and related practices. Product gains and skill-building are not the same outcome.
La IA hizo la escritura más rápida y mejorCambios en tiempo de escritura y calidad evaluada independientemente en un experimento prerregistradocon 444 personas.Made for Hugo L. Alvarez, by HanademiFuentes: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.; Graham, S., & Perin,D. (2007). Writing next: Effective strategies to improve writing of adolescents in middle and high schools. Carnegie Corporation of New York.Las fuentes usan bases de medición distintas; lea la comparación como tendencia, no como una escala exacta.Unidadcambio porcentualReducción de tiempo40 %Aumento de calidad18 %
ChatGPT redujo un 40% el tiempo de escritura profesional y elevó un 18% la calidad evaluada independientemente. Sin embargo, evidencia anterior a la IA encontró un efecto estandarizado de 0,82 asociado con enseñanza explícita de estrategias y prácticas relacionadas. Mejorar el producto y desarrollar habilidades no son el mismo resultado.
A 293-person experiment found betterstories but less diversity.The study measured creative-writing outcomes, not long-term development ofindependent writers.Sources: Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. ScienceAdvances, 10(28).
AI can improve one writer's product while making many writers more alike. In a 293-person experiment, access to AI ideas improved individual story evaluations but reduced similarity-adjusted diversity across writers. Universities should value original reasoning, not only polished output.
Un experimento con 293 personas hallómejores relatos, pero menor diversidad.El estudio midió resultados de escritura creativa, no el desarrollo a largo plazo deescritores independientes.Fuentes: Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. ScienceAdvances, 10(28).
La IA puede mejorar el producto de un autor mientras hace que muchos autores se parezcan más. En un experimento con 293 personas, acceder a ideas de IA mejoró las evaluaciones individuales de relatos, pero redujo la diversidad ajustada por similitud. Las universidades deben valorar el razonamiento original, no solo un producto pulido.
AI users reported gains in both speed andqualityPost-use responses from an Australian Government Microsoft 365 Copilot trial.Made for Hugo L. Alvarez, by HanademiSources: digital.gov.au.Unitshare of respondents69%Faster completion61%Improved quality
Users experience AI as useful because its immediate benefits are visible. In the Australian Government trial, 69% reported faster completion and 61% reported improved work quality. Educational designs that ask students to struggle first must compete with that reward.
Los usuarios de IA reportaron mejoras enrapidez y calidadRespuestas posteriores al uso en una prueba de Microsoft 365 Copilot del Gobierno australiano.Made for Hugo L. Alvarez, by HanademiFuentes: digital.gov.au.Unidadproporción de encuestados69 %Finalización más rápida61 %Mejor calidad
Los usuarios perciben la IA como útil porque sus beneficios inmediatos son visibles. En la prueba del Gobierno australiano, el 69% reportó terminar más rápido y el 61% señaló una mejor calidad. Los diseños educativos que piden al estudiante esforzarse primero deben competir con esa recompensa.
Legal models hallucinated in 58% to88% of tested queries.The reported range is methodology-dependent and applies to the tested legal tasks, notevery AI response.Sources: Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal ofLegal Analysis.; Profiling Legal Hallucinations in Large Language Models.
Fluency is not evidence. Three general-purpose language models hallucinated in roughly 58% to 88% of tested legal queries, depending on model and task. Students need explicit habits for checking sources and claims.
Los modelos jurídicos alucinaron enentre el 58% y el 88%.El rango reportado depende de la metodología y se aplica a las tareas jurídicasevaluadas, no a toda respuesta de IA.Fuentes: Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal ofLegal Analysis.; Profiling Legal Hallucinations in Large Language Models.
La fluidez no es evidencia. Tres modelos lingüísticos generales alucinaron en aproximadamente entre el 58% y el 88% de las consultas jurídicas evaluadas, según el modelo y la tarea. Los estudiantes necesitan hábitos explícitos para comprobar fuentes y afirmaciones.
Creative-thinking gaps existed beforesustained ChatGPT exposurePISA 2022 creative-thinking points for the OECD average and Singapore.Made for Hugo L. Alvarez, by HanademiSources: OECD. (2024). PISA 2022 results, Volume III: Creative minds, creative schools. OECD Publishing.UnitPISA points01020304033OECD average41SingaporeOECD average
The incoming cohort did not start from one uniform skill level. In PISA 2022, Singapore scored 41 in creative thinking while the OECD average was 33. These results predate sustained school exposure to ChatGPT.
Las brechas de pensamiento creativo existían antesde la exposición sostenida a ChatGPTPuntos de pensamiento creativo de PISA 2022 para el promedio de la OCDE y Singapur.Made for Hugo L. Alvarez, by HanademiFuentes: OECD. (2024). PISA 2022 results, Volume III: Creative minds, creative schools. OECD Publishing.Unidadpuntos PISA01020304033Promedio OCDE41SingaporePromedio OCDE
La cohorte que ingresa no partió de un nivel uniforme de habilidades. En PISA 2022, Singapur obtuvo 41 puntos en pensamiento creativo y el promedio de la OCDE fue 33. Estos resultados son anteriores a la exposición escolar sostenida a ChatGPT.
Digital distraction already reached mostmathematics classroomsPISA 2022 students reporting device distraction in at least some mathematics lessons.Made for Hugo L. Alvarez, by HanademiSources: OECD. (2023). PISA 2022 results, Volume II: Learning during and from disruption. OECD Publishing.Unitshare of students65%Own device59%Classmates' devices
Independent learning requires sustained attention, but digital distraction predates widespread generative AI. PISA 2022 found 65% of students were distracted by their own devices and 59% by classmates' devices in at least some mathematics lessons. AI adds immediate answers to an already contested attention environment.
La distracción digital ya alcanzaba a lamayoría en clases de matemáticasEstudiantes de PISA 2022 que reportaron distracción por dispositivos en al menos algunas clases dematemáticas.Made for Hugo L. Alvarez, by HanademiFuentes: OECD. (2023). PISA 2022 results, Volume II: Learning during and from disruption. OECD Publishing.Unidadproporción de estudiantes65 %Dispositivo propio59 %Dispositivos de compañeros
El aprendizaje independiente exige atención sostenida, pero la distracción digital es anterior a la difusión de la IA generativa. PISA 2022 encontró que el 65% se distraía con sus propios dispositivos y el 59% con los de sus compañeros en al menos algunas clases de matemáticas. La IA añade respuestas inmediatas a un entorno de atención ya disputado.
ACT readiness was falling before generativeAI spreadAverage United States ACT composite score, 2020 and 2024.Made for Hugo L. Alvarez, by HanademiSources: ACT. (2024). The graduating class of 2024 national profile report.UnitACT composite score19.520.020.520202024The average reached 19.4 in2024.20.6
The average ACT composite fell from 20.6 in 2020 to 19.4 in 2024. The share meeting all four readiness benchmarks also declined. Timing alone cannot assign this longer-running deterioration to generative AI.
La preparación medida por ACT ya caíaantes de difundirse la IA generativaPuntuación compuesta promedio de ACT en Estados Unidos, 2020 y 2024.Made for Hugo L. Alvarez, by HanademiFuentes: ACT. (2024). The graduating class of 2024 national profile report.Unidadpuntuación compuesta ACT19,520,020,520202024El promedio llegó a 19,4 en 2024.20,6
El promedio compuesto de ACT cayó de 20,6 en 2020 a 19,4 en 2024. La proporción que cumplía los cuatro umbrales de preparación también disminuyó. Las fechas por sí solas no permiten atribuir este deterioro previo a la IA generativa.
Readiness fell across ACT, reading, and mathematicsbefore the 2026 cohort arrivedMade for Hugo L. Alvarez, by HanademiSources: ACT. (2024). The graduating class of 2024 national profile report.; 2024 ACT National Graduating ClassExecutive Summary.; National Center for Education Statistics. (2025). The Nation's Report Card: 2024mathematics and reading assessments.UnitIndex, starting score equals 100ACT composite94.2-5.8%100NAEP mathematics, grade 896.5-3.5%100NAEP reading, grade 898.1-1.9%100949698100Base 100; ACT: 2020 to 2024; NAEP: 2019 to 2024
The ACT composite declined 5.8%, eighth-grade mathematics 3.5%, and eighth-grade reading 1.9%. The broad decline predates the first nearly fully AI-exposed class and complicates causal attribution.
La preparación cayó en ACT, lectura y matemáticasantes de la llegada de la generación de 2026Made for Hugo L. Alvarez, by HanademiFuentes: ACT. (2024). The graduating class of 2024 national profile report.; 2024 ACT National Graduating ClassExecutive Summary.; National Center for Education Statistics. (2025). The Nation's Report Card: 2024mathematics and reading assessments.UnidadÍndice, la puntuación inicial equivale a 100Resultado compuesto de ACT94,2-5,8%100Matemáticas de NAEP, octavo grado96,5-3,5%100Lectura de NAEP, octavo grado98,1-1,9%100949698100Base 100; ACT: 2020 a 2024; NAEP: 2019 a 2024
El resultado compuesto de ACT cayó un 5,8%, las matemáticas de octavo grado un 3,5% y la lectura de octavo grado un 1,9%. El descenso general precede a la primera generación casi totalmente expuesta a la IA y complica la atribución causal.
Eighth-grade reading declined before andafter ChatGPT launchedNational eighth-grade reading average, NAEP scale points, 2019 to 2024.Made for Hugo L. Alvarez, by HanademiSources: National Center for Education Statistics. (2025). The Nation's Report Card: 2024 mathematics and reading assessments.UnitNAEP scale points258260262264201920222024Reading reached 258 in 2024.
National eighth-grade reading fell from 263 in 2019 to 259 in 2022 and 258 in 2024. The decline began before public ChatGPT and continued afterward. The incoming cohort therefore carries a preexisting literacy shock alongside AI exposure.
La lectura de octavo grado cayó antes ydespués del lanzamiento de ChatGPTPromedio nacional de lectura de octavo grado, puntos de la escala NAEP, 2019 a 2024.Made for Hugo L. Alvarez, by HanademiFuentes: National Center for Education Statistics. (2025). The Nation's Report Card: 2024 mathematics and readingassessments.Unidadpuntos de escala NAEP258260262264201920222024La lectura llegó a 258 en 2024.
La lectura nacional de octavo grado bajó de 263 en 2019 a 259 en 2022 y 258 en 2024. La caída comenzó antes de ChatGPT público y continuó después. Por tanto, la cohorte que ingresa carga con una alteración previa de alfabetización junto con la exposición a la IA.
The reported mathematics decline remainsprovisionalReported national eighth-grade mathematics average, NAEP scale points, 2019 to 2024. The values werenot independently reverified at build time.Made for Hugo L. Alvarez, by HanademiSources: National Center for Education Statistics. (2025). The Nation's Report Card: 2024 mathematics and reading assessments.UnitNAEP scale points270275280201920222024The reported 2024 average is272.282
The supplied figures report a fall from 282 in 2019 to 274 in 2022 and 272 in 2024. That pattern matches the broader preexisting-readiness warning. Because these values were not independently reverified at build time, the slide frames them as provisional rather than settled evidence.
La caída reportada en matemáticas siguesiendo provisionalPromedio nacional reportado de matemáticas de octavo grado, puntos de la escala NAEP, 2019 a 2024. Losvalores no se volvieron a verificar de forma independiente.Made for Hugo L. Alvarez, by HanademiFuentes: National Center for Education Statistics. (2025). The Nation's Report Card: 2024 mathematics and readingassessments.Unidadpuntos de escala NAEP270275280201920222024El promedio reportado de 2024 es272.282
Las cifras suministradas reportan una caída de 282 en 2019 a 274 en 2022 y 272 en 2024. Ese patrón coincide con la advertencia más amplia sobre preparación previa. Como los valores no se volvieron a verificar de forma independiente, se presentan como provisionales y no como evidencia definitiva.
OpenAI's detector caught 26% and falselyflagged 9%OpenAI classifier performance on AI-written and human-written text before withdrawal in 2023.Made for Hugo L. Alvarez, by HanademiSources: OpenAI. (2023). New AI classifier for indicating AI-written text.Unitclassification rateAI text detected26%Human text misclassified9%
OpenAI's own classifier failed in both directions. It correctly identified only 26% of AI-written text and mislabeled 9% of human-written text. OpenAI withdrew it, leaving universities without a reliable shortcut for proving authorship.
El detector de OpenAI identificó el 26% yacusó falsamente al 9%Rendimiento del clasificador de OpenAI sobre texto de IA y texto humano antes de su retirada en 2023.Made for Hugo L. Alvarez, by HanademiFuentes: OpenAI. (2023). New AI classifier for indicating AI-written text.Unidadtasa de clasificaciónTexto de IA detectado26 %Texto humano mal clasificado9 %
El propio clasificador de OpenAI falló en ambas direcciones. Identificó correctamente solo el 26% del texto escrito por IA y clasificó erróneamente como IA el 9% del texto humano. OpenAI lo retiró, dejando a las universidades sin un atajo confiable para demostrar la autoría.
OpenAI's classifier missed 2.8 AI texts forevery one it caughtMade for Hugo L. Alvarez, by HanademiSources: OpenAI. (2023). New AI classifier for indicating AI-written text.; New AI classifier for indicating AI-writtentext.; Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native Englishwriters. Patterns.UnitShare of evaluated textsAI text, OpenAI classifierCaught 26%Missed 74%Human text, OpenAI classifierCleared 91%False flag 9%Non-native essays: overallCleared 38.8%False flag 61.2%Non-native essays: any detectorCleared 3%Flagged 97%050100Share of texts
The classifier missed most AI writing while still falsely flagging human work. In the non-native essay evaluation, 61.2% were classified as AI-generated and 97% were flagged by at least one detector.
El clasificador de OpenAI omitió 2,8 textosde IA por cada uno que detectóMade for Hugo L. Alvarez, by HanademiFuentes: OpenAI. (2023). New AI classifier for indicating AI-written text.; New AI classifier for indicatingAI-written text.; Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased againstnon-native English writers. Patterns.UnidadProporción de textos evaluadosTextos de IA, clasificador OpenAIDetectó 26%Omitió 74%Textos humanos, clasificador OpenAINo marcado 91%Error 9%Ensayos no nativos: totalNo marcado 38,8%Marcado por error 61,2%Ensayos no nativos: algún detectorNo marcado 3%Marcado 97%050100Proporción de textos
El clasificador omitió la mayoría de los textos de IA y aun así marcó erróneamente trabajos humanos. En la evaluación de ensayos de autores no nativos, el 61,2% fue clasificado como generado por IA y el 97% fue marcado por al menos un detector.
One detector study found severe risk fornon-native writersReported detector flags on 91 TOEFL essays by non-native English writers. The values were notindependently reverified at build time.Made for Hugo L. Alvarez, by HanademiSources: Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers.Patterns.Unitshare of essaysAverage AI classification61.2%Flagged by any detector97%
The danger is not evenly distributed. In one study, detectors classified 61.2% of 91 TOEFL essays by non-native English writers as AI-generated, and 97% were flagged by at least one detector. Because the values were not independently reverified here, the finding is presented as a study-specific warning.
Un estudio de detectores halló un riesgograve para autores no nativosAlertas reportadas de detectores sobre 91 ensayos TOEFL de autores no nativos de inglés. Los valores nose volvieron a verificar de forma independiente.Made for Hugo L. Alvarez, by HanademiFuentes: Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native Englishwriters. Patterns.Unidadproporción de ensayosClasificación promedio como IA61,2 %Señalado por algún detector97 %
El peligro no se distribuye de forma uniforme. En un estudio, los detectores clasificaron como IA el 61,2% de 91 ensayos TOEFL de autores no nativos de inglés, y el 97% fue señalado por al menos un detector. Como los valores no se volvieron a verificar aquí, el hallazgo se presenta como una advertencia específica del estudio.
Student AI use outpaced knowledge andinstitutional supportResponses from an international student survey in 2024.Made for Hugo L. Alvarez, by HanademiSources: Digital Education Council. (2024). Global AI student survey 2024.; Diliberti, M. K., Schwartz, H. L., Doan, S., Shapiro, A.,Rainey, L. R., & Lake, R. J. (2024). Using artificial intelligence tools in K-12 classrooms. RAND Corporation.Sources use different measurement bases; read the comparison directionally, not as one exact scale.Unitshare of studentsRegular AI use86%Insufficient knowledge58%Inadequate support80%
Usage is not the same as literacy. The international survey reported 86% regular AI use, yet 58% felt insufficiently knowledgeable and 80% judged institutional support inadequate. In United States schools, RAND separately reported 25% teacher use during 2023-24 and substantially higher use among principals.
El uso estudiantil de IA superó alconocimiento y al apoyo institucionalRespuestas de una encuesta internacional a estudiantes en 2024.Made for Hugo L. Alvarez, by HanademiFuentes: Digital Education Council. (2024). Global AI student survey 2024.; Diliberti, M. K., Schwartz, H. L., Doan, S.,Shapiro, A., Rainey, L. R., & Lake, R. J. (2024). Using artificial intelligence tools in K-12 classrooms. RAND Corporation.Las fuentes usan bases de medición distintas; lea la comparación como tendencia, no como una escala exacta.Unidadproporción de estudiantesUso regular de IA86 %Conocimiento insuficiente58 %Apoyo inadecuado80 %
El uso no equivale a alfabetización. La encuesta internacional reportó un 86% de uso regular de IA, pero el 58% se consideraba poco preparado y el 80% juzgaba insuficiente el apoyo institucional. En escuelas estadounidenses, RAND reportó por separado un 25% de uso docente durante 2023-24 y un uso mucho mayor entre directores.
Regular AI use ran 66 points ahead ofadequate institutional supportMade for Hugo L. Alvarez, by HanademiSources: Digital Education Council. (2024). Global AI student survey 2024.; Survey: 86% of Students Already UseAI in Their Studies.; UNESCO. (2023). UNESCO survey: Less than 10% of schools and universities have formalguidance on AI.UnitShare of surveyed students or institutionsRegular student AI use86%Felt knowledgeable42%44 pp86%Rated support adequate20%66 pp86%Institutions with formal guidance<10%>76 pp86%050100Share
The student survey yields gaps of 44 points for knowledge and 66 points for adequate support. A separate institutional survey puts formal guidance more than 76 points below regular student use.
El uso habitual de IA superó por 66 puntosal apoyo institucional adecuadoMade for Hugo L. Alvarez, by HanademiFuentes: Digital Education Council. (2024). Global AI student survey 2024.; Survey: 86% of Students Already UseAI in Their Studies.; UNESCO. (2023). UNESCO survey: Less than 10% of schools and universities have formalguidance on AI.UnidadProporción de estudiantes o instituciones encuestadosUso habitual de IA por estudiantes86%Se sentían informados42%44 p.p.86%Consideraban adecuado el apoyo20%66 p.p.86%Instituciones con orientación formal<10%>76 p.p.86%050100Proporción
La encuesta estudiantil muestra brechas de 44 puntos en conocimiento y 66 puntos en apoyo adecuado. Otra encuesta institucional sitúa la orientación formal más de 76 puntos por debajo del uso habitual de los estudiantes.
The rules remain unsettledA hand writes an AI-use choice on a chalkboard listing several options.Sources: Confusing school policies on AI, ChatGPT use leave families guessing. axios.com.
Unclear choices about whether and when AI is allowed widen the gap between student use and institutional support.
Las reglas siguen sin definirseUna mano escribe una opción sobre el uso de la IA en una pizarra que enumera varias alternativas.Fuentes: Confusing school policies on AI, ChatGPT use leave families guessing. axios.com.
Las decisiones poco claras sobre si se permite la IA y cuándo amplían la brecha entre el uso estudiantil y el apoyo institucional.
The strongest response is to teach thinkingexplicitlyEstimated additional months of progress under favorable implementation in EEF evidence syntheses.Made for Hugo L. Alvarez, by HanademiSources: Education Endowment Foundation. (2025). Teaching and Learning Toolkit.Unitadditional months of progress8Metacognition7Reading strategies6Feedback
The answer is not to recreate a world without AI. Evidence syntheses estimate 8 additional months of progress from metacognition, 7 from reading-comprehension strategies, and 6 from feedback under favorable implementation. Universities should make these processes visible in assessment.
La mejor respuesta es enseñar a pensar deforma explícitaMeses adicionales estimados de progreso con una implementación favorable en las síntesis de evidencia deEEF.Made for Hugo L. Alvarez, by HanademiFuentes: Education Endowment Foundation. (2025). Teaching and Learning Toolkit.Unidadmeses adicionales de progreso8Metacognición7Estrategias de lectura6Retroalimentación
La respuesta no es recrear un mundo sin IA. Las síntesis de evidencia estiman 8 meses adicionales de progreso por metacognición, 7 por estrategias de comprensión lectora y 6 por retroalimentación con una implementación favorable. Las universidades deben hacer visibles estos procesos en la evaluación.
The reported retrieval advantage remainsprovisionalOne-week recall after repeated retrieval and repeated study. The values were not independently reverifiedat build time.Made for Hugo L. Alvarez, by HanademiSources: Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-termretention. Psychological Science.Unitone-week recallRepeated retrieval61%Repeated study40%
The supplied study reports 61% recall after repeated retrieval and 40% after repeated study one week later. Repeated study performed better immediately, which helps explain why easier methods feel effective. Because the values were not independently reverified here, the finding remains provisional.
La ventaja reportada de recuperación siguesiendo provisionalRecuerdo después de una semana tras recuperación repetida y estudio repetido. Los valores no se volvierona verificar de forma independiente.Made for Hugo L. Alvarez, by HanademiFuentes: Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improveslong-term retention. Psychological Science.Unidadrecuerdo tras una semanaRecuperación repetida61 %Estudio repetido40 %
El estudio suministrado reporta un 61% de recuerdo tras recuperación repetida y un 40% tras estudio repetido una semana después. El estudio repetido rindió mejor de inmediato, lo que ayuda a explicar por qué los métodos más fáciles parecen eficaces. Como los valores no se volvieron a verificar aquí, el hallazgo sigue siendo provisional.
Completion stayed high while readiness stillweakenedGross lower-secondary completion rates in the United Kingdom and United States, percent of the relevantage group.Made for Hugo L. Alvarez, by HanademiSources: World Bank. (n.d.). Lower secondary completion rate, total (% of relevant age group) [SE.SEC.CMPT.LO.ZS]. WorldBank Open Data.Unitgross completion rate98%100%102%105%2016201820202022UnitedKingdomUnited States
The United Kingdom and United States maintained high gross lower-secondary completion rates. Yet the United States readiness measures in this deck still declined. Finishing school and mastering the skills universities expect are different outcomes.
La finalización siguió alta mientras lapreparación se debilitabaTasas brutas de finalización de secundaria básica en Reino Unido y Estados Unidos, porcentaje del grupo deedad correspondiente.Made for Hugo L. Alvarez, by HanademiFuentes: World Bank. (n.d.). Lower secondary completion rate, total (% of relevant age group) [SE.SEC.CMPT.LO.ZS]. WorldBank Open Data.Unidadtasa bruta de finalización98 %100 %102 %105 %2016201820202022ReinoUnidoEstados Unidos
Reino Unido y Estados Unidos mantuvieron tasas brutas altas de finalización de secundaria básica. Sin embargo, las medidas estadounidenses de preparación incluidas en esta presentación disminuyeron. Terminar la escuela y dominar las habilidades que esperan las universidades son resultados distintos.
Completion data can move sharply withoutmeasuring learningGross lower-secondary completion rates in Singapore and Türkiye, percent of the relevant age group.Made for Hugo L. Alvarez, by HanademiSources: World Bank. (n.d.). Lower secondary completion rate, total (% of relevant age group) [SE.SEC.CMPT.LO.ZS]. WorldBank Open Data.Unitgross completion rate100%120%2010201220142016201820202022SingaporeTürkiyeTürkiye's gross rate reached131% in 2020.
Singapore's gross rate stayed near or above 100% across much of the series. Türkiye's rate spiked to 131% in 2020 before falling to 91.8% in 2023. Gross completion measures system participation and cohort structure, not whether students can reason or write independently.
Los datos de finalización pueden variarmucho sin medir aprendizajeTasas brutas de finalización de secundaria básica en Singapur y Türkiye, porcentaje del grupo de edadcorrespondiente.Made for Hugo L. Alvarez, by HanademiFuentes: World Bank. (n.d.). Lower secondary completion rate, total (% of relevant age group) [SE.SEC.CMPT.LO.ZS]. WorldBank Open Data.Unidadtasa bruta de finalización100 %120 %2010201220142016201820202022SingapurTürkiyeLa tasa bruta de Türkiye llegó al131% en 2020.
La tasa bruta de Singapur se mantuvo cerca o por encima del 100% durante gran parte de la serie. La tasa de Türkiye subió al 131% en 2020 antes de caer al 91,8% en 2023. La finalización bruta mide participación y estructura de cohortes, no si los estudiantes pueden razonar o escribir independientemente.
Assess what students can explain, verify,and reproduce without assistance.The evidence points to a design test: AI can improve output, but assessment must keeplearning visible and independently testable.Sources: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.; Generative AI without guardrailscan harm learning: Evidence from high school mathematics.; Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.;Experimental Evidence on the Productivity Effects of ....; OpenAI. (2023). New AI classifier for indicating AI-written text.; New AI classifier for indicating AI-written text.
AI can improve the submitted product while weakening the evidence that the student learned. Its design determines whether it supports reasoning or replaces it. Universities should observe the process, require verification, and test independent performance.
Hay que evaluar lo que el estudiante puedeexplicar, verificar y reproducir sin ayuda.La evidencia apunta a una prueba de diseño: la IA puede mejorar el resultado, pero laevaluación debe mantener el aprendizaje visible y comprobable de forma independiente.Fuentes: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.; Generative AI without guardrailscan harm learning: Evidence from high school mathematics.; Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.;Experimental Evidence on the Productivity Effects of ....; OpenAI. (2023). New AI classifier for indicating AI-written text.; New AI classifier for indicating AI-written text.
La IA puede mejorar el producto entregado mientras debilita la evidencia de que el estudiante aprendió. Su diseño determina si apoya el razonamiento o lo sustituye. Las universidades deben observar el proceso, exigir verificación y evaluar el desempeño independiente.
Crossing GPT-4's competence frontierreversed the performance advantageMade for Hugo L. Alvarez, by HanademiSources: Dell'Acqua, F., McFowland, E., Mollick, E. R., et al. (2023). Navigating the jagged technological frontier. Harvard Business School WorkingPaper 24-013.; ChatGPT: A Harvard Business School study of Boston Consulting Group workers found the AI tool significantly improved theirperformance.; Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.UnitReported percentage change or percentage-point changeWithin tested competenceMore tasks completed+12.2%Faster+25.1%Higher-rated output>40%Outside tested competenceCorrect answer-19 pp-2002040Reported change
Inside the frontier, consultants completed more work faster and at higher quality. Outside it, AI assistance reduced the likelihood of a correct answer by 19 points.
Cruzar la frontera de competencia deGPT-4 revirtió la ventaja de rendimientoMade for Hugo L. Alvarez, by HanademiFuentes: Dell'Acqua, F., McFowland, E., Mollick, E. R., et al. (2023). Navigating the jagged technological frontier. Harvard Business School WorkingPaper 24-013.; ChatGPT: A Harvard Business School study of Boston Consulting Group workers found the AI tool significantly improved theirperformance.; Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.UnidadCambio porcentual o cambio en puntos porcentuales informadoDentro de la competencia evaluadaMás tareas completadas+12,2%Mayor rapidez+25,1%Mayor calidad evaluada>40%Fuera de la competencia evaluadaAcierto-19 p.p.-2002040Cambio informado
Dentro de la frontera, los consultores completaron más trabajo, con mayor rapidez y calidad. Fuera de ella, la asistencia de IA redujo en 19 puntos la probabilidad de una respuesta correcta.
In summaryMade for Hugo L. Alvarez, by HanademiSources: OpenAI. (2023). New AI classifier for indicating AI-written text.; Bastani, H., et al. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.; Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E.(2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis.; Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harmlearning. NBER Working Paper No. 33097.; Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.; Sidoti, O., & Gottfried, J. (2025). About a quarter…OpenAI withdrew its classifier after reporting 26% correct detection of AI-written text and 9% false classification of human text as AI-written.After AI removal, unrestricted-GPT students scored 17% below controls, while safeguarded-tutor students showed nostatistically significant unaided disadvantage.Three general-purpose language models hallucinated in approximately 58% to 88% of tested legal queries, depending on model and task.In a Turkish high-school mathematics experiment, unrestricted GPT raised practice performance 48%, while asafeguarded AI tutor raised it 127% relative to control.On a task outside GPT-4's competence frontier, assisted BCG consultants were 19 percentage points less likely to producethe correct answer than controls.Pew reported 26% of United States teens used ChatGPT for schoolwork in 2024, up from 13% in 2023, with demographic differences.
The value of the research is not only what each source knew, but what became visible when their evidence was combined.
En resumenMade for Hugo L. Alvarez, by HanademiFuentes: OpenAI. (2023). New AI classifier for indicating AI-written text.; Bastani, H., et al. (2024). Generative AI can harm learning. NBER Working Paper No. 33097.; Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E.(2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis.; Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harmlearning. NBER Working Paper No. 33097.; Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.; Sidoti, O., & Gottfried, J. (2025). About a quarter…OpenAI retiró su clasificador tras informar un 26% de detección correcta de texto de IA y un 9% de falsos positivos sobre texto humano.Tras retirar la IA, los estudiantes con GPT sin restricciones puntuaron un 17% por debajo del control, mientras el tutorprotegido no mostró una desventaja significativa.Tres modelos lingüísticos generales alucinaron en aproximadamente entre el 58% y el 88% de las consultas jurídicas, según modelo y tarea.En un experimento turco de matemáticas de secundaria, GPT sin restricciones elevó la práctica un 48% y un tutorprotegido un 127% frente al control.En una tarea fuera de la frontera de competencia de GPT-4, los consultores asistidos tuvieron 19 puntos porcentualesmenos de probabilidad de acertar que el control.Pew informó que el 26% de los adolescentes estadounidenses usó ChatGPT para tareas en 2024, frente al 13% en 2023,con diferencias demográficas.
El valor de la investigación no está solo en cada fuente, sino en lo que apareció al combinar sus evidencias.

The research behind this deck

The design of the assistance mattered more than simply providing AI. This cohort encountered public AI near the start of high school.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

El diseño de la ayuda importó más que el simple acceso a la IA. Esta cohorte encontró la IA pública cerca del inicio de secundaria.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English

Related researchInvestigación relacionada