La muestra cayó de 54 en las sesiones 1-3 a 18 en la sesión 4.
Start with the split: 96% on the take-home midterm and 48.6% on the controlled final. Professor Roberto Serrano also said he had conclusive evidence against at least 50 students in a class of 86. This is not causal proof, but it shows how quickly assisted output can separate from independently demonstrated mastery.
Empecemos por la división: 96% en el parcial en casa y 48,6% en el final controlado. El profesor Roberto Serrano también afirmó tener pruebas concluyentes contra al menos 50 estudiantes de un curso de 86. No es una prueba causal, pero muestra la rapidez con que el resultado asistido puede separarse del dominio demostrado sin ayuda.
This is the human face of the Brown dispute. Serrano’s claim was not about occasional misuse, but evidence involving at least 50 students. The larger question is whether traditional take-home assessment can still distinguish assistance from mastery.
Este es el rostro humano de la disputa de Brown. La afirmación de Serrano no trataba de un uso indebido ocasional, sino de pruebas relacionadas con al menos 50 estudiantes. La pregunta mayor es si la evaluación tradicional en casa todavía puede distinguir la asistencia del dominio.
This is larger than one course or one professor. Across 500,000 grades, courses most exposed to writing and coding gained 13 percentage points in A grades after ChatGPT. Average GPA rose only 0.12, making the concentration of the shift the important clue.
Esto es mayor que un curso o un profesor. En 500.000 calificaciones, los cursos más expuestos a escritura y programación ganaron 13 puntos porcentuales de calificaciones A después de ChatGPT. El promedio general subió solo 0,12, por lo que la concentración del cambio es la pista importante.
Adoption did not creep forward. In one year, assessment use rose 35 points and any AI use rose 26 points. Among students using any AI, the supplied analysis estimates that 95.7% used generative AI for assessments in 2025.
La adopción no avanzó lentamente. En un año, el uso en evaluaciones subió 35 puntos y el uso de cualquier IA aumentó 26. Entre quienes usaban alguna IA, el análisis suministrado estima que el 95,7% empleó IA generativa en evaluaciones en 2025.
This is the cleanest picture of the output-mastery split. Homework scores rose 18% and completion time fell 30%, both attractive outcomes. Yet monthly examinations fell 20% and entrance examinations declined between 18% and 24%.
Esta es la imagen más clara de la división entre resultado y dominio. Las notas de las tareas subieron un 18% y el tiempo de finalización cayó un 30%, dos resultados atractivos. Sin embargo, los exámenes mensuales cayeron un 20% y los de ingreso disminuyeron entre un 18% y un 24%.
Learning time fell most where AI was easiest to use. High school students reduced time by 31.3% and university students by 26.9%. The danger is not exposure to AI itself, but repeatedly removing the effort that practice is meant to create.
El tiempo de aprendizaje cayó más donde la IA era más fácil de usar. Los estudiantes de secundaria superior redujeron el tiempo un 31,3% y los universitarios un 26,9%. El peligro no es la exposición a la IA en sí, sino eliminar repetidamente el esfuerzo que la práctica debe producir.
The immediate convenience was followed by a measurable retention penalty. On randomly assigned and proctored questions, the chance of a correct answer declined cumulatively by 25%. This is a decline in probability, not a 25-point fall in examination scores.
A la comodidad inmediata le siguió una penalización medible en la retención. En preguntas asignadas al azar y supervisadas, la probabilidad de acertar cayó de forma acumulada un 25%. Es una caída de probabilidad, no una disminución de 25 puntos en las notas.
The difference appears after the immediate task is over. Traditional learners retained 68.5% of the knowledge, while AI-assisted learners retained 57.5%. On-demand support can feel like easier training while leaving less strength behind.
La diferencia aparece después de terminar la tarea inmediata. Quienes estudiaron de forma tradicional retuvieron el 68,5% del conocimiento, mientras quienes usaron IA retuvieron el 57,5%. La ayuda bajo demanda puede hacer que el entrenamiento parezca más fácil y dejar menos fortaleza después.
The model was the same kind of tool; the protocol changed. Unrestricted help produced a 30% gain, while controlled timing produced 64%. The useful AI behaves like a coach who chooses the moment to intervene, not a substitute available for every move.
El tipo de herramienta era el mismo; cambió el protocolo. La ayuda sin restricciones produjo un avance del 30%, mientras el control del momento produjo un 64%. La IA útil actúa como un entrenador que elige cuándo intervenir, no como un sustituto disponible para cada movimiento.
The evidence is not uniformly negative. A 2026 review of 34 K-12 studies found a moderate positive overall effect on cognitive learning. The distinction matters: these were purpose-built educational agents, not unrestricted chatbots completing homework.
La evidencia no es uniformemente negativa. Una revisión de 2026 con 34 estudios escolares halló un efecto general positivo moderado sobre el aprendizaje cognitivo. La distinción importa: eran agentes educativos diseñados para ese fin, no chatbots sin restricciones que hacían las tareas.
The volume of research is much larger than the dependable causal core. Stanford found only 20 papers with strong cause-and-effect evidence in a repository containing more than 800. The field may be promising, but it is still radically immature.
El volumen de investigación es mucho mayor que su núcleo causal fiable. Stanford encontró solo 20 trabajos con evidencia sólida de causa y efecto en un repositorio con más de 800. El campo puede ser prometedor, pero todavía es radicalmente inmaduro.
The strongest hopeful result combines AI with teachers rather than removing them. A six-week program in Nigeria produced gains estimated as 1.5 to 2 years of ordinary schooling. That figure is a randomized cost-effectiveness comparison, not the literal time students attended.
El resultado esperanzador más fuerte combina la IA con docentes en lugar de eliminarlos. Un programa de seis semanas en Nigeria produjo avances estimados como 1,5 a 2 años de escolaridad habitual. Esa cifra es una comparación aleatoria de rentabilidad educativa, no el tiempo literal de asistencia.
These studies preserve a role for both the teacher and the learner. Tutor CoPilot raised mastery by 9 points for students with lower-rated tutors. In a separate trial, AI feedback increased immediate correction by 16% and next-problem success by 7%.
Estos estudios preservan un papel para el docente y el estudiante. Tutor CoPilot elevó el dominio en 9 puntos entre alumnos con tutores de menor valoración. En otro ensayo, la retroalimentación con IA aumentó la corrección inmediata un 16% y el éxito en el problema siguiente un 7%.
The learning gains came with a reliability problem. Initial help failed quality checks on 32% of problems. Added safeguards reduced failures to zero in the algebra set and 13% in statistics, showing that deployment quality depends on checking.
Los avances de aprendizaje llegaron con un problema de fiabilidad. La ayuda inicial falló los controles en el 32% de los problemas. Las salvaguardas redujeron los fallos a cero en álgebra y al 13% en estadística, lo que muestra que la calidad del despliegue depende de la verificación.
Education faces a two-sided reliability problem. Raw AI math help failed on 32% of problems, while detector errors reached 31% for Originality and 39% for Turnitin.
La educación enfrenta un problema de fiabilidad por ambos lados. La ayuda matemática sin filtrar falló en el 32% de los problemas, mientras que los errores de los detectores alcanzaron el 31% en Originality y el 39% en Turnitin.
The studies point to a common protocol. The student tries first, receives limited feedback, revises, and must transfer the learning to another answer. Teachers and safeguards remain in the loop because both timing and reliability change the outcome.
Los estudios apuntan a un protocolo común. El estudiante intenta primero, recibe retroalimentación limitada, revisa y debe transferir el aprendizaje a otra respuesta. Los docentes y las salvaguardas permanecen dentro del proceso porque tanto el momento como la fiabilidad cambian el resultado.
AI did not create the baseline crisis. From 2018 to 2022, OECD mathematics performance fell 15 points, a decline already equal to 18.1% of the 83-point spread between Singapore and Ireland in the listed PISA ranking.
La IA no creó la crisis de partida. Entre 2018 y 2022, el rendimiento de matemáticas de la OCDE cayó 15 puntos, una disminución equivalente al 18,1% de la brecha de 83 puntos entre Singapur e Irlanda en la clasificación PISA mostrada.
The benchmark is demanding. Singapore reached 575 points, followed by Macao China at 552 and Chinese Taipei at 547. Any claim that AI will transform education must be judged against systems already producing exceptional measured mastery.
La referencia es exigente. Singapur alcanzó 575 puntos, seguido de Macao, China, con 552 y Taipéi Chino con 547. Cualquier afirmación de que la IA transformará la educación debe juzgarse frente a sistemas que ya producen un dominio medido excepcional.
Technology access alone has a weak educational record. A 10-year randomized follow-up in 531 rural Peruvian schools found better computer skills but no significant academic gains. A separate, contested study of 4,600 schools found phone pouches had almost no overall effect on test scores.
El acceso a la tecnología por sí solo tiene un historial educativo débil. Un seguimiento aleatorio de 10 años en 531 escuelas rurales peruanas halló mejores habilidades informáticas, pero no avances académicos significativos. Otro estudio disputado de 4.600 escuelas encontró que las bolsas para móviles casi no afectaron las notas generales.
The debate can become abstract, but the stakes are students. They inherit a learning decline that began before public generative AI. The opportunity is to use new tools to repair learning systems rather than repeat the assumption that access alone creates achievement.
El debate puede volverse abstracto, pero lo que está en juego son los estudiantes. Heredan una caída del aprendizaje que comenzó antes de la IA generativa pública. La oportunidad consiste en usar nuevas herramientas para reparar los sistemas educativos y no repetir la suposición de que el acceso por sí solo produce rendimiento.
The psychological labor-market shock is already here. Forty-one percent said AI changed the jobs they wanted, and 61% expected fewer opportunities across many fields. Yet only 48% felt prepared, making career guidance an immediate educational responsibility.
El impacto psicológico en el mercado laboral ya está aquí. El 41% dijo que la IA cambió los empleos que quería y el 61% esperaba menos oportunidades en muchos sectores. Sin embargo, solo el 48% se sentía preparado, lo que convierte la orientación profesional en una responsabilidad educativa inmediata.
Capability and deployment are running on different clocks. Eight in ten senior executives reported no effect on employment or productivity in their organizations. That does not predict a quiet future, but it explains why youth anxiety can lead measured workplace change.
La capacidad y el despliegue avanzan con relojes distintos. Ocho de cada diez altos directivos no informaron efectos sobre el empleo o la productividad en sus organizaciones. Eso no predice un futuro tranquilo, pero explica por qué la ansiedad juvenil puede adelantarse al cambio medido en el trabajo.
In 2023, AI systems solved 4.4% of SWE-bench problems. One year later, they solved 71.7%. Schools cannot make assessment durable simply by selecting tasks that today’s models still fail.
En 2023, los sistemas de IA resolvieron el 4,4% de los problemas de SWE-bench. Un año después resolvieron el 71,7%. Las escuelas no pueden hacer duradera una evaluación simplemente eligiendo tareas que los modelos actuales todavía no resuelven.
Benchmark scores are only one sign of movement. Since 2019, the length of software tasks completed with 50% reliability has doubled roughly every seven months. Curriculum and assessment therefore need designs that survive changing capability, not a frozen list of AI-proof tasks.
Las puntuaciones de las pruebas son solo una señal del movimiento. Desde 2019, la duración de las tareas de software completadas con un 50% de fiabilidad se ha duplicado aproximadamente cada siete meses. El currículo y la evaluación necesitan diseños que sobrevivan a capacidades cambiantes, no una lista congelada de tareas resistentes a la IA.
The forecast is not simple mass replacement. It projects 170 million jobs created and 92 million displaced, leaving a net gain of 78 million. Youth anxiety remains rational because the number displaced is larger than the entire net gain.
La previsión no describe un reemplazo masivo simple. Proyecta 170 millones de empleos creados y 92 millones desplazados, con un aumento neto de 78 millones. La ansiedad juvenil sigue siendo racional porque el número desplazado supera toda la ganancia neta.
The internet offers a caution against treating technological adoption as permanent mass unemployment. Use rose from 43.1% in 2000 to 94.7% in 2024, while youth unemployment moved through crises and reached 9.34% in 2025, close to 9.28% in 2000. This does not prove AI will follow the same path, but it shows why adaptation and transition matter more than a static job count.
Internet ofrece una advertencia contra interpretar la adopción tecnológica como desempleo masivo permanente. El uso subió del 43,1% en 2000 al 94,7% en 2024, mientras el desempleo juvenil atravesó varias crisis y llegó al 9,34% en 2025, cerca del 9,28% de 2000. Esto no demuestra que la IA seguirá el mismo camino, pero muestra por qué la adaptación y la transición importan más que una cifra laboral estática.
The research behind this deck
The submitted work looked excellent, but controlled performance and the professor’s allegations told a different story. A single disputed course became a warning about whether grades still measure independent knowledge.
Key findings
Among 26,811 Chinese students, AI raised homework scores 18% and cut time 30%, yet monthly exams fell 20% and entrance exams 18-24% over two years. Centre for Economic Policy Research
In one Brown economics class, the take-home midterm averaged 96%, but the controlled final averaged 48.6%, amid the professor's allegations of widespread AI cheating. The Register
Brown professor Roberto Serrano said he had conclusive evidence that at least 50 students cheated on the March midterm in an 86-student class. EL PAÍS
Across 500,000 university grades, writing- and coding-heavy courses gained 13 percentage points in A grades after ChatGPT, while average GPA rose 0.12. University of California, Berkeley
A global reanalysis found 91.1% of 22,963 university students used ChatGPT, but only 39.1% believed it improved their critical thinking. Mendeley Data
After ChatGPT, time spent on AI-friendly math problems fell 31.3% for high schoolers, 26.9% for college students and 9.0% for middle schoolers. University of California, Irvine and McGraw Hill
In a randomized trial, traditionally studying students retained 68.5% of knowledge long term, compared with 57.5% among students using AI for answers or study help. Social Sciences & Humanities Open
The influential MIT brain-activity study involved only 54 participants, and just 18 completed its final session, so it cannot establish lasting brain damage. MIT Media Lab
A separate 2026 review of 34 K-12 studies also found a moderate positive overall effect from AI agents on cognitive learning outcomes. Elsevier
The argument
Across 500,000 university grades, A grades rose 13 points in writing- and coding-heavy courses.
Assessment use almost caught total AI use, leaving institutions little time to redesign evaluation.
AI made homework faster and better while performance without the same assistance deteriorated.
The easier a problem was to outsource, the less time older students spent working through it.
The cost of reduced practice appeared later, when students had to retrieve the knowledge under supervision.
AI assistance preserved 57.5% of knowledge, compared with 68.5% under traditional study.
This research is published in English and Spanish. Ver en español