Hanademi

AI can produce answers without producing learningLa IA puede producir respuestas sin producir aprendizaje · V7

29 slides · 26 min · 2026-08-04 language
ENES
theme
LightDark
view
DeckTableTalk
brand
HanademiPlatzi
AI can produce answerswithout producing learningMade for Eduardo Santiago, by Hanademi
  1. Fluent performance is not dependable competence.
  2. AI raised undergraduate knowledge by 0.27 SD, with gains lasting one week.
  3. Removing struggle can improve the worksheet while weakening the learner.
La IA puede producirrespuestas sin produciraprendizajeMade for Eduardo Santiago, by Hanademi
  1. Un rendimiento fluido no equivale a una competencia fiable.
  2. La IA elevó 0,27 DE el conocimiento universitario y la mejora duró una semana.
  3. Eliminar el esfuerzo puede mejorar la tarea y debilitar al estudiante.
AI helps until the task changesChanges in consulting outcomes with GPT-4 versus controls, 2023 field experiment.Made for Eduardo Santiago, by HanademiSources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effectsof AI on knowledge worker productivity and quality. Harvard Business School Working Paper 24-013.The final value is a percentage-point accuracy difference.UnitChange versus controlSuitable tasks12.2%Speed25.1%Output quality40%Outside-frontier accuracy-19%
The same experiment produced a spectacular win and a serious failure. GPT-4 improved speed, completion, and quality on suitable consulting tasks. On a task outside its capability frontier, assisted consultants lost 19 points in accuracy.
La IA ayuda hasta que cambia la tareaCambios en resultados de consultoría con GPT-4 frente al control, experimento de campo de 2023.Made for Eduardo Santiago, by HanademiFuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effectsof AI on knowledge worker productivity and quality. Harvard Business School Working Paper 24-013.El valor final es una diferencia de precisión en puntos porcentuales.UnidadCambio frente al controlTareas adecuadas12,2 %Rapidez25,1 %Calidad del resultado40 %Precisión fuera de frontera-19 %
El mismo experimento produjo un éxito espectacular y un fracaso serio. GPT-4 mejoró la rapidez, la finalización y la calidad en tareas de consultoría adecuadas. En una tarea fuera de su frontera de capacidad, los consultores con IA perdieron 19 puntos de precisión.
A few terms before we beginMade for Eduardo Santiago, by HanademiStandard deviationA common way to compare learning effects across different tests.Capability frontierThe boundary between tasks AI handles well and poorly.Unaided transferUsing learned knowledge on a new task without AI assistance.Cognitive offloadingLetting an external tool remember or solve part of a task.
These terms separate visible performance from durable learning. Effect sizes allow comparisons across different tests. Unaided transfer asks whether students can still reason when the tool is gone.
Algunos términos antes de comenzarMade for Eduardo Santiago, by HanademiDesviación estándarUna forma común de comparar efectos de aprendizaje entre pruebasdiferentes.Frontera de capacidadEl límite entre tareas que la IA resuelve bien y mal.Transferencia sin ayudaUsar lo aprendido en una tarea nueva sin ayuda de IA.Descarga cognitivaDejar que una herramienta externa recuerde o resuelva parte de una tarea.
Estos términos separan el rendimiento visible del aprendizaje duradero. Los tamaños de efecto permiten comparar pruebas distintas. La transferencia sin ayuda pregunta si los estudiantes todavía pueden razonar cuando la herramienta desaparece.
Higher-order thinking begins beyond recallSix ordered cognitive processes in the revised Bloom taxonomy, from remembering to creating.Made for Eduardo Santiago, by HanademiThe values encode the six-level order, not measured distances between cognitive processes.UnitCognitive level6Create5Evaluate4Analyze3Apply2Understand1Remember
Higher-order thinking is not answer production. The revised taxonomy distinguishes 6 processes, culminating in analysis, evaluation, and creation. Knowing a recipe is different from diagnosing why the cake failed and redesigning it.
El pensamiento de orden superior comienzamás allá del recuerdoSeis procesos cognitivos ordenados en la taxonomía revisada de Bloom, desde recordar hasta crear.Made for Eduardo Santiago, by HanademiLos valores codifican el orden de seis niveles, no distancias medidas entre procesos cognitivos.UnidadNivel cognitivo6Crear5Evaluar4Analizar3Aplicar2Comprender1Recordar
El pensamiento de orden superior no consiste en producir respuestas. La taxonomía revisada distingue 6 procesos y culmina en analizar, evaluar y crear. Conocer una receta es distinto de diagnosticar por qué falló el pastel y rediseñarla.
Deeper thinking improves with deliberatepracticeStandardized effects from separate studies of self-explanation, critical-thinking instruction, and policydebate.Made for Eduardo Santiago, by HanademiSources: Abrami, P. C., et al. (2008). Instructional interventions affecting critical thinking skills and dispositions: A stage 1 meta-analysis. Review of Educational Research.; Bisra, K.,Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30, 703-725.; edworkingpapers.com.These studies cover different populations and interventions.UnitStandard deviationsPrompted self-explanation0.6Critical-thinking instruction0.3Boston policy debate0.1
Critical thinking is teachable, but it does not appear automatically. Interventions across 117 studies averaged an effect near 0.30 standard deviations. Prompted self-explanation reached about 0.55, while Boston policy debate produced analytical gains of 0.13.
El pensamiento profundo mejora conpráctica deliberadaEfectos estandarizados de estudios separados sobre autoexplicación, enseñanza de pensamiento crítico ydebate de políticas.Made for Eduardo Santiago, by HanademiFuentes: Abrami, P. C., et al. (2008). Instructional interventions affecting critical thinking skills and dispositions: A stage 1 meta-analysis. Review of Educational Research.; Bisra,K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30, 703-725.; edworkingpapers.com.Estos estudios cubren poblaciones e intervenciones distintas.UnidadDesviaciones estándarAutoexplicación guiada0,6Enseñanza de pensamientocrítico0,3Debate de políticas en Boston0,1
El pensamiento crítico se puede enseñar, pero no aparece automáticamente. Las intervenciones en 117 estudios promediaron un efecto cercano a 0,30 desviaciones estándar. La autoexplicación guiada alcanzó cerca de 0,55, mientras el debate de políticas en Boston produjo mejoras analíticas de 0,13.
Digital tools are already in the classroomStudents use tablets and laptops while working individually in a classroom.Sources: UNESCO. (2023). UNESCO survey: Less than 10% of schools and universities have formal guidance on AI.
Because devices already mediate classroom work, deeper thinking must be deliberately built into each task.
Las herramientas digitales ya están en el aulaEstudiantes usan tabletas y computadoras portátiles mientras trabajan individualmente en un aula.Fuentes: UNESCO. (2023). UNESCO survey: Less than 10% of schools and universities have formal guidance on AI.
Como los dispositivos ya median el trabajo en el aula, el pensamiento profundo debe incorporarse deliberadamente en cada tarea.
Student AI use was already routine in 2024Share of 3,839 surveyed students in 16 countries using AI regularly, weekly, or daily in 2024.Made for Eduardo Santiago, by HanademiSources: Digital Education Council. (2024). Global AI student survey 2024.UnitStudents86%Regular users54%Weekly users24%Daily users
Student adoption moved quickly. In this 2024 survey, 86% used AI regularly, 54% used it weekly, and 24% used it daily. The survey covered 3,839 students across 16 countries.
El uso estudiantil de IA ya era habitual en2024Proporción de 3.839 estudiantes encuestados en 16 países que usaba IA regularmente, cada semana o adiario en 2024.Made for Eduardo Santiago, by HanademiFuentes: Digital Education Council. (2024). Global AI student survey 2024.UnidadEstudiantes86 %Usuarios habituales54 %Usuarios semanales24 %Usuarios diarios
La adopción estudiantil avanzó rápidamente. En esta encuesta de 2024, el 86% usaba IA regularmente, el 54% cada semana y el 24% a diario. La encuesta cubrió a 3.839 estudiantes de 16 países.
The student AI-use share is over 8.6 timesthe institutional guidance shareDivide each student-survey percentage by the 10% upper bound for institutions with formal guidance; the trueratios are larger because guidance was below 10%.Made for Eduardo Santiago, by HanademiSources: Digital Education Council. (2024). Global AI student survey 2024.; Student Use of Generative AI inHigher Education: Patterns, Gaps, and Institutional Readiness.; UNESCO. (2023). UNESCO survey: Less than 10% ofschools and universities have formal guidance on AI.Unitminimum ratio to formal-guidance share, xRegular AI use8.6University integration falls short8Want more AI training7.2Feel underinformed5.8
Even the share feeling underinformed is at least 5.8 times the share of institutions with formal guidance.
La proporción de uso estudiantil de IA supera en 8,6veces la proporción de instituciones con directricesDividir cada porcentaje de la encuesta estudiantil por el límite superior de 10% para instituciones condirectrices formales; las razones reales son mayores porque la cifra estaba por debajo de 10%.Made for Eduardo Santiago, by HanademiFuentes: Digital Education Council. (2024). Global AI student survey 2024.; Student Use of Generative AI in HigherEducation: Patterns, Gaps, and Institutional Readiness.; UNESCO. (2023). UNESCO survey: Less than 10% ofschools and universities have formal guidance on AI.Unidadrazón mínima frente a la proporción con directrices formales, xUso habitual de IA8,6Integración universitaria insuficiente8Quieren más formación en IA7,2Se sienten poco informados5,8
Incluso la proporción que se siente poco informada es al menos 5,8 veces la proporción de instituciones con directrices formales.
AI assistance reaches the workbookA phone with a voice-assistant interface is held beside an open workbook.Sources: AI Use in Schools Is Quickly Increasing but Guidance Lags Behind. rand.org.
When AI assistance sits beside the assignment, guidance must reach students at the point of use.
La asistencia de IA llega al cuadernoUn teléfono con una interfaz de asistente de voz se sostiene junto a un cuaderno abierto.Fuentes: AI Use in Schools Is Quickly Increasing but Guidance Lags Behind. rand.org.
Cuando la asistencia de IA está junto a la tarea, la orientación debe llegar al alumnado en el momento de uso.
ChatGPT made professional writing 40%fasterChanges in time, evaluator-rated quality, and completion among 453 professionals in a randomized 2023experiment.Made for Eduardo Santiago, by HanademiSources: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence.Science.UnitChange versus controlTime reduction40%Quality increase18%Completion increase10%
This is the strongest case for AI assistance. Professionals finished 40% faster, improved quality by 18%, and completed 10% more work. But the experiment measured task output, not durable learning.
ChatGPT aceleró un 40% la escrituraprofesionalCambios en tiempo, calidad evaluada y finalización entre 453 profesionales en un experimento aleatorizadode 2023.Made for Eduardo Santiago, by HanademiFuentes: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence.Science.UnidadCambio frente al controlReducción de tiempo40 %Aumento de calidad18 %Aumento de finalización10 %
Este es el caso más sólido a favor de la asistencia con IA. Los profesionales terminaron un 40% antes, mejoraron la calidad un 18% y completaron un 10% más de trabajo. Pero el experimento midió el resultado de la tarea, no el aprendizaje duradero.
AI raised undergraduate knowledge by0.27 SD, with gains lasting one week.This is an important counterpoint. AI can support learning under some designs, but theclaim bank does not provide a numeric delayed effect.Sources: germanr.com.
Not every AI-assisted gain disappears when the session ends. In one randomized undergraduate experiment, AI access raised immediate knowledge scores by 0.27 standard deviations. The improvement persisted one week later.
La IA elevó 0,27 DE el conocimientouniversitario y la mejora duró una semana.Este es un contrapunto importante. La IA puede apoyar el aprendizaje con algunosdiseños, pero el banco de evidencia no aporta un efecto retrasado numérico.Fuentes: germanr.com.
No toda mejora asistida por IA desaparece al terminar la sesión. En un experimento aleatorizado con universitarios, el acceso a IA elevó 0,27 desviaciones estándar los resultados inmediatos de conocimiento. La mejora persistió una semana.
Answer-giving AI helped practice, then hurtthe examRelative performance in secondary mathematics with unrestricted GPT, safeguarded GPT tutoring, and alater unaided exam.Made for Eduardo Santiago, by HanademiSources: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. Universityof Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.The assisted and unaided results are distinct phases of the same secondary-mathematics intervention.UnitRelative performanceUnrestricted practice48%Safeguarded practice127%Later unaided exam-17%
During practice, both AI conditions produced large gains. The safeguarded tutor performed especially well because it supported reasoning rather than freely supplying answers. After AI removal, the unrestricted group scored 17% below controls.
La IA que daba respuestas ayudó en lapráctica y perjudicó el examenRendimiento relativo en matemáticas de secundaria con GPT sin restricciones, tutor GPT con salvaguardas yun examen posterior sin ayuda.Made for Eduardo Santiago, by HanademiFuentes: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. University ofPennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.Los resultados asistidos y sin ayuda son fases distintas de la misma intervención de matemáticas en secundaria.UnidadRendimiento relativoPráctica sin restricciones48 %Práctica con salvaguardas127 %Examen posterior sin ayuda-17 %
Durante la práctica, ambas condiciones con IA produjeron grandes mejoras. El tutor con salvaguardas funcionó especialmente bien porque apoyaba el razonamiento en vez de entregar respuestas libremente. Tras retirar la IA, el grupo sin restricciones obtuvo un 17% menos que el control.
Homework rose as exam scores fellChange after staggered AI adoption among 26,811 Chinese students, including outcomes observed within sixmonths.Made for Eduardo Santiago, by HanademiSources: studylib.net.UnitChangeHomework score18%Completion time-30%Monthly exam score-20%
The pattern repeats at much larger scale. Homework scores increased and completion time fell after AI adoption. Yet monthly exam scores declined 20% within six months.
Las tareas mejoraron mientras cayeron losexámenesCambio tras la adopción escalonada de IA entre 26.811 estudiantes chinos, incluidos resultados observadosen seis meses.Made for Eduardo Santiago, by HanademiFuentes: studylib.net.UnidadCambioNota de tareas18 %Tiempo de finalización-30 %Nota de examen mensual-20 %
El patrón se repite a una escala mucho mayor. Las notas de tareas aumentaron y el tiempo de finalización disminuyó tras adoptar IA. Sin embargo, las notas de exámenes mensuales bajaron un 20% en seis meses.
The tool is not the learningA tablet is held in front of a stylized AI chip and brain display.Sources: AI homework tools cut exam scores by 20%, study of 26,000 Chinese students finds | The Star. thestar.com.my.
AI can sit at the center of a task while independent performance remains the test of learning.
La herramienta no es el aprendizajeUna tableta se sostiene frente a una pantalla estilizada con un chip de IA y un cerebro.Fuentes: AI homework tools cut exam scores by 20%, study of 26,000 Chinese students finds | The Star. thestar.com.my.
La IA puede ocupar el centro de una tarea, pero el desempeño independiente sigue siendo la prueba del aprendizaje.
AI helped most where tutors neededsupportTopic-mastery increase from Tutor CoPilot across about 900 tutors and 1,800 students in authenticsessions.Made for Eduardo Santiago, by HanademiSources: Wang, R. E., et al. (2024). Tutor CoPilot: A human-AI approach for scaling real-time expertise. NBER Working Paper33097.UnitPercentage-point gainAll students4%Lower-rated tutors9%2.2x
Tutor CoPilot did not replace tutors. It gave about 900 tutors real-time support across sessions involving 1,800 students. Mastery rose 4 points overall and 9 points for students of lower-rated tutors.
La IA ayudó más donde los tutoresnecesitaban apoyoAumento del dominio temático con Tutor CoPilot entre unos 900 tutores y 1.800 estudiantes en sesionesauténticas.Made for Eduardo Santiago, by HanademiFuentes: Wang, R. E., et al. (2024). Tutor CoPilot: A human-AI approach for scaling real-time expertise. NBERWorking Paper 33097.UnidadMejora en puntos porcentualesTodos los estudiantes4 %Tutores con menor calificación9 %2,2x
Tutor CoPilot no sustituyó a los tutores. Dio apoyo en tiempo real a unos 900 tutores en sesiones con 1.800 estudiantes. El dominio aumentó 4 puntos en general y 9 puntos entre estudiantes de tutores con menor calificación.
In one 2026 assessment, AI tutoringscored 4.3Estimated scores transcribed from a published figure: pre-score, active in-class lesson, and AI lesson.Made for Eduardo Santiago, by HanademiSources: AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authenticeducational setting | Scientific Reports. (2026). Published figure transcribed from pixels.Values were estimated from bar heights to approximately 0.1 score.UnitAssessment score4.3AI lesson3.5Active in-class lesson2.7Pre-score
This assessment points toward genuine instructional promise. The AI lesson reached an estimated post-score of 4.3, compared with 3.5 for the active in-class lesson. The values were read from a published image, so the precision is limited.
En una evaluación de 2026, la tutoría conIA obtuvo 4,3Puntuaciones estimadas y transcritas de una figura publicada: puntuación previa, lección activa presencial ylección con IA.Made for Eduardo Santiago, by HanademiFuentes: AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in anauthentic educational setting | Scientific Reports. (2026). Published figure transcribed from pixels.Los valores se estimaron a partir de la altura de las barras con precisión aproximada de 0,1.UnidadPuntuación de evaluación4,3Lección con IA3,5Lección activa presencial2,7Puntuación previa
Esta evaluación apunta a una promesa educativa real. La lección con IA alcanzó una puntuación posterior estimada de 4,3, frente a 3,5 para la lección activa presencial. Los valores se leyeron de una imagen publicada, por lo que la precisión es limitada.
Modern tutoring falls far short of 2 sigmaStandardized effects from a modern tutoring meta-analysis and scaled online mathematics tutoring in Spain.Made for Eduardo Santiago, by HanademiSources: Nickow, A., Oreopoulos, P., & Quan, V. (2024). The impressive effects of tutoring on preK-12 learning: A systematic review andmeta-analysis. Review of Educational Research.; Bloom, B. S. (1984). The 2 sigma problem. Educational Researcher.; edworkingpapers.com.The 2-sigma reference is Bloom's historical benchmark, not the average result of modern tutoring.UnitStandard deviations0120.4Tutoring meta-analysis0.1Spain course grades0.1Spain standardized testsBloom benchmark: 2 sigma
Tutoring has a meaningful average effect near 0.37 standard deviations. When Spain scaled online mathematics tutoring, effects reached 0.15 for grades and 0.11 for standardized tests. Those results remain far below Bloom's famous 2-sigma benchmark.
La tutoría moderna queda muy lejos de 2sigmaEfectos estandarizados de un metaanálisis moderno de tutoría y de tutoría virtual de matemáticas ampliadaen España.Made for Eduardo Santiago, by HanademiFuentes: Nickow, A., Oreopoulos, P., & Quan, V. (2024). The impressive effects of tutoring on preK-12 learning: A systematic review andmeta-analysis. Review of Educational Research.; Bloom, B. S. (1984). The 2 sigma problem. Educational Researcher.; edworkingpapers.com.La referencia de 2 sigma es el punto histórico de Bloom, no el resultado promedio de la tutoría moderna.UnidadDesviaciones estándar0120,4Metaanálisis de tutoría0,1Notas de curso en España0,1Pruebas estandarizadas enEspañaReferencia de Bloom: 2 sigma
La tutoría tiene un efecto promedio significativo cercano a 0,37 desviaciones estándar. Cuando España amplió la tutoría virtual de matemáticas, los efectos llegaron a 0,15 en notas y 0,11 en pruebas estandarizadas. Esos resultados siguen muy por debajo de la famosa referencia de 2 sigma de Bloom.
Active learning improves examperformanceSource-provided summary values from 225 undergraduate STEM studies, including the 0.47 SD effect.Made for Eduardo Santiago, by HanademiSources: Freeman, S., et al. (2014). Active learning increases student performance in science, engineering, andmathematics. Proceedings of the National Academy of Sciences.The third mark is 225 studies divided by 100, exactly as supplied.UnitSource-provided display valuesLecture baseline0Active-learning effect0.5Studies divided by 1002.2
Across 225 undergraduate STEM studies, active learning improved examination performance by 0.47 standard deviations. That result is stronger evidence than passive answer consumption. AI should create more attempts, explanations, and revisions, not remove them.
El aprendizaje activo mejora el rendimientoen exámenesValores resumidos aportados por la fuente de 225 estudios universitarios STEM, incluido el efecto de 0,47DE.Made for Eduardo Santiago, by HanademiFuentes: Freeman, S., et al. (2014). Active learning increases student performance in science, engineering, andmathematics. Proceedings of the National Academy of Sciences.La tercera marca representa 225 estudios divididos entre 100, exactamente como se suministró.UnidadValores de presentación de la fuenteBase de clase magistral0Efecto de aprendizaje activo0,5Estudios divididos entre 1002,2
En 225 estudios universitarios STEM, el aprendizaje activo mejoró 0,47 desviaciones estándar el rendimiento en exámenes. Ese resultado tiene más respaldo que el consumo pasivo de respuestas. La IA debe crear más intentos, explicaciones y revisiones, no eliminarlos.
Confident AI references can be badly wrongFabricated-reference rates in one bibliographic study and hallucination rates in separate legal tests.Made for Eduardo Santiago, by HanademiSources: Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports,13, 14045.; Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models.Journal of Legal Analysis.UnitPercentFabricated references55%18%GPT-3.5GPT-4Legal hallucinations58%69%72%88%GPT-4ChatGPTPaLM 2Llama 2
An AI answer can sound like a confident witness with an unreliable memory. One study found GPT-3.5 fabricated 55% of references and GPT-4 fabricated 18%. Separate legal tests reported hallucination rates from 58% to 88%.
Las referencias seguras de la IA puedenestar muy equivocadasTasas de referencias inventadas en un estudio bibliográfico y tasas de alucinación en pruebas jurídicasseparadas.Made for Eduardo Santiago, by HanademiFuentes: Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. ScientificReports, 13, 14045.; Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in largelanguage models. Journal of Legal Analysis.UnidadPorcentajeReferencias inventadas55 %18 %GPT-3.5GPT-4Alucinaciones jurídicas58 %69 %72 %88 %GPT-4ChatGPTPaLM 2Llama 2
Una respuesta de IA puede sonar como un testigo seguro con memoria poco fiable. Un estudio encontró que GPT-3.5 inventó el 55% de las referencias y GPT-4 el 18%. Pruebas jurídicas separadas informaron tasas de alucinación entre el 58% y el 88%.
Verification demands a closer lookAn eye is seen through a torn opening in paper.Sources: Frontiers | Amplifier or substitute? A systematic review of generative AI’s impact on higher-order cognitive skills among university students. frontiersin.org.
Students must look past fluent output and inspect the evidence underneath.
La verificación exige mirar más de cercaUn ojo se ve a través de una abertura rasgada en papel.Fuentes: Frontiers | Amplifier or substitute? A systematic review of generative AI’s impact on higher-order cognitive skills among university students. frontiersin.org.
El alumnado debe mirar más allá de una respuesta fluida y examinar la evidencia que la sustenta.
Human TOEFL essays triggered AIdetectorsDetection outcomes across seven tools for 91 human-written TOEFL essays, 2023 study.Made for Eduardo Santiago, by HanademiSources: Liang, W., et al. (2023). GPT detectors are biased against non-native English writers. Patterns.UnitHuman essays flaggedAverage detector61.2%At least one detector97%1.6x
These essays were written by people, yet detectors repeatedly called them machine-generated. Across seven tools, 61.2% were flagged on average. At least one detector flagged 97% of the 91 essays.
Ensayos TOEFL humanos activarondetectores de IAResultados de siete herramientas para 91 ensayos TOEFL escritos por humanos, estudio de 2023.Made for Eduardo Santiago, by HanademiFuentes: Liang, W., et al. (2023). GPT detectors are biased against non-native English writers. Patterns.UnidadEnsayos humanos señaladosDetector promedio61,2 %Al menos un detector97 %1,6x
Estos ensayos fueron escritos por personas, pero los detectores los clasificaron repetidamente como generados por máquinas. En siete herramientas, el 61,2% fue señalado en promedio. Al menos un detector señaló el 97% de los 91 ensayos.
Assessment should gather evidence ofprocess, authorship, and explanation.These safeguards produce multiple forms of evidence. None requires treating aprobabilistic detector as a final verdict.Sources: Tertiary Education Quality and Standards Agency. (2023). Assessment reform for the age of artificial intelligence.Equal values indicate inclusion in the four-part safeguard set, not measured effectiveness.
A safer misconduct process does not begin and end with a detector score. Version histories reveal the work's development. Oral questions, supervised baseline work, and source verification test whether the student understands and owns the submission.
La evaluación debe reunir evidencia delproceso, la autoría y la explicación.Estas salvaguardas producen varias formas de evidencia. Ninguna exige tratar undetector probabilístico como veredicto final.Fuentes: Tertiary Education Quality and Standards Agency. (2023). Assessment reform for the age of artificial intelligence.Los valores iguales indican inclusión en el conjunto de cuatro salvaguardas, no eficacia medida.
Un proceso más seguro ante una posible falta no empieza ni termina con una puntuación de detector. Los historiales de versiones revelan el desarrollo del trabajo. Las preguntas orales, el trabajo base supervisado y la verificación de fuentes comprueban si el estudiante comprende y domina la entrega.
AI lifted weaker performers moreQuality gains inside GPT-4's capability frontier for below-average and above-average consultants.Made for Eduardo Santiago, by HanademiSources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.UnitQuality gain43%Below-average consultants17%Above-average consultants
Inside the capability frontier, below-average consultants gained about 43% in quality. Above-average consultants gained about 17%. AI can therefore compress visible performance gaps during assisted work.
La IA impulsó más a quienes rendíanmenosMejoras de calidad dentro de la frontera de GPT-4 para consultores por debajo y por encima del promedio.Made for Eduardo Santiago, by HanademiFuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper24-013.UnidadMejora de calidad43 %Consultores bajo el promedio17 %Consultores sobre el promedio
Dentro de la frontera de capacidad, los consultores por debajo del promedio ganaron cerca del 43% en calidad. Los consultores por encima ganaron alrededor del 17%. La IA puede reducir las brechas visibles durante el trabajo asistido.
Targeted support and guardrails magnify AIgains by 2.25x to 2.65xDivide 127 by 48 for safeguarded versus unrestricted GPT, 43 by 17 for below-average versusabove-average consultants, and 9 by 4 for the lower-rated tutor subgroup versus the overall result.Made for Eduardo Santiago, by HanademiSources: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. Universityof Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.;Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.Unitgain multiplier, xSafeguarded vs unrestricted GPT2.6Below-average vs above-average consultants2.5Lower-rated tutor subgroup vs overall2.2
Three distinct comparisons cluster around a 2.5x gain multiplier when AI is constrained or directed toward a weaker baseline.
El apoyo dirigido y las salvaguardas multiplican lasmejoras con IA entre 2,25x y 2,65xDividir 127 por 48 para GPT con salvaguardas frente a GPT sin restricciones, 43 por 17 para consultorespor debajo y por encima del promedio, y 9 por 4 para el subgrupo de tutores con menor valoración frente alresultado general.Made for Eduardo Santiago, by HanademiFuentes: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. Universityof Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.;Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.Unidadmúltiplo de la mejora, xGPT con salvaguardas frente a GPT sin restricciones2,6Consultores por debajo frente a por encima del promedio2,5Subgrupo de tutores con menor valoración frente al total2,2
Tres comparaciones distintas se concentran alrededor de un múltiplo de mejora de 2,5x cuando la IA está limitada o dirigida a un punto de partida más débil.
Assess both the finished work and thestudent's command of it.The framework is evidence-informed guidance rather than a conclusively validateduniversal protocol.Sources: Tertiary Education Quality and Standards Agency. (2023). Assessment reform for the age of artificial intelligence.; UNESCO. (2023). Guidancefor generative AI in education and research. UNESCO.Equal areas show five included layers, not measured weights or proven effectiveness.
The submitted product is no longer complete evidence of learning. A stronger design combines 5 layers: unaided baseline work, a record of AI use, verification, reflection, and oral defense. Each layer answers a different question about capability and authorship.
Evalúe tanto el trabajo final como el dominioque tiene el estudiante.El marco es una orientación informada por evidencia, no un protocolo universal validadoconcluyentemente.Fuentes: Tertiary Education Quality and Standards Agency. (2023). Assessment reform for the age of artificial intelligence.; UNESCO. (2023). Guidancefor generative AI in education and research. UNESCO.Las áreas iguales muestran cinco capas incluidas, no pesos medidos ni eficacia demostrada.
El producto entregado ya no es una prueba completa del aprendizaje. Un diseño más sólido combina 5 capas: trabajo base sin ayuda, registro del uso de IA, verificación, reflexión y defensa oral. Cada capa responde una pregunta distinta sobre capacidad y autoría.
A course succeeds only when learningsurvives the removal of assistance.This is the decisive redesign. AI collaboration can remain legitimate, but independentcapability must stay visible across five separate outcomes.Sources: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.; Bastani, H., et al.(2024). Generative AI can harm learning. University of Pennsylvania working paper.Each bar marks one distinct outcome to measure separately, not equal importance or effect size.
Course evaluation must stop treating one polished product as the whole outcome. Measure assisted output, then test unaided transfer and delayed retention. Add source accuracy and explanation quality to reveal whether students can defend what they submitted.
Un curso solo tiene éxito cuando elaprendizaje sobrevive al retirar la ayuda.Este es el rediseño decisivo. La colaboración con IA puede seguir siendo legítima, perola capacidad independiente debe permanecer visible en cinco resultados separados.Fuentes: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science.; Bastani, H., et al.(2024). Generative AI can harm learning. University of Pennsylvania working paper.Cada barra marca un resultado distinto que debe medirse por separado, no igual importancia ni tamaño de efecto.
La evaluación del curso debe dejar de tratar un producto pulido como el resultado completo. Mida el resultado asistido y después pruebe la transferencia sin ayuda y la retención retrasada. Añada precisión de fuentes y calidad de explicación para saber si los estudiantes pueden defender lo entregado.
Design for what students can explain,verify, and repeat unaided.The goal is not to ban assistance. It is to make assistance strengthen reasoning ratherthan replace it.Sources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.; Jagged Frontier – Why AI Helps on Some Tasks and Hurts on Others | AI wiki.;Bastani, H., et al. (2024). Generative AI can harm learning. University of Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.; Liang, W., et al.(2023). GPT detectors are biased against non-native English writers. Patterns.; Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13…
AI can raise output while hiding weaker understanding. Guardrails, active practice, verification, and human tutoring change that equation. The decisive test is what students can explain and transfer after assistance is removed.
Diseñe para lo que los estudiantes puedanexplicar, verificar y repetir sin ayuda.El objetivo no es prohibir la ayuda. Es lograr que la ayuda fortalezca el razonamiento envez de sustituirlo.Fuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.; Jagged Frontier – Why AI Helps on Some Tasks and Hurts on Others | AI wiki.;Bastani, H., et al. (2024). Generative AI can harm learning. University of Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high school mathematics.; Liang, W., et al.(2023). GPT detectors are biased against non-native English writers. Patterns.; Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13…
La IA puede elevar el resultado mientras oculta una comprensión más débil. Las salvaguardas, la práctica activa, la verificación y la tutoría humana cambian esa ecuación. La prueba decisiva es lo que los estudiantes pueden explicar y transferir después de retirar la ayuda.
Seven-detector screening flagged about 88of 91 human essaysMultiply 91 essays by the reported 97%, 61.2%, and below-1% rates; the Turnitin value is an upper boundunder its threshold condition.Made for Eduardo Santiago, by HanademiSources: Liang, W., et al. (2023). GPT detectors are biased against non-native English writers. Patterns.; Turnitin.(2023). AI writing detection model: False positive rate and testing methodology.; Turnitin's AI detector:higher-than-expected false positives.Unithuman essays flagged per 91At least one of seven88.3Average detector55.7Turnitin threshold condition0.9
At least one detector flagged about 88 human essays, while the average detector flagged about 56.
El cribado con siete detectores marcó cercade 88 de 91 ensayos humanosMultiplicar 91 ensayos por las tasas reportadas de 97%, 61,2% y menos de 1%; el valor de Turnitin es unlímite superior bajo la condición de su umbral.Made for Eduardo Santiago, by HanademiFuentes: Liang, W., et al. (2023). GPT detectors are biased against non-native English writers. Patterns.; Turnitin.(2023). AI writing detection model: False positive rate and testing methodology.; Turnitin's AI detector:higher-than-expected false positives.Unidadensayos humanos marcados por cada 91Al menos uno de siete88,3Detector promedio55,7Condición del umbral de Turnitin0,9
Al menos un detector marcó cerca de 88 ensayos humanos, mientras que el detector promedio marcó unos 56.
AI output gains can reverse on independentor shifted tasksMade for Eduardo Santiago, by HanademiSources: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning.University of Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high schoolmathematics.; Bastani, H., et al. (2024). Generative AI can harm learning. University of Pennsylvania working paper.Unitreported percent change or percentage-point changeIndependent or shiftedAssisted or in frontierMath tutoring−17+48Consulting work−19+40China study−20+18−200+20+40Reported change
The positive-to-negative swing spans 65 points in mathematics, 59 in consulting, and 38 in the China study.
Las mejoras con IA pueden revertirse entareas independientes o desplazadasMade for Eduardo Santiago, by HanademiFuentes: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning.University of Pennsylvania working paper.; Generative AI without guardrails can harm learning: Evidence from high schoolmathematics.; Bastani, H., et al. (2024). Generative AI can harm learning. University of Pennsylvania working paper.Unidadcambio porcentual o en puntos porcentuales reportadoIndependiente o desplazadaAsistida o dentro de fronteraTutoría matemática−17+48Trabajo de consultoría−19+40Estudio en China−20+18−200+20+40Cambio reportado
El giro de positivo a negativo alcanza 65 puntos en matemáticas, 59 en consultoría y 38 en el estudio de China.
In summaryMade for Eduardo Santiago, by HanademiSources: Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045.; Freeman, S., et al. (2014). Active learning increases studentperformance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences.; Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis.Educational Psychology Review, 30, 703-725.; Wang, R. E., et al. (2024). Tutor CoPilot: A human-AI approach for scaling real-time expertise. NBER Working Paper 33097.; Abrami, P. C., et al. (2008). Instructional…One bibliographic study found fabricated-reference rates of 55% for GPT-3.5 and 18% for GPT-4.Across 225 undergraduate STEM studies, active learning improved examination performance by 0.47 standard deviations.A meta-analysis of 64 studies estimated that prompted self-explanation improved learning by approximately 0.55 standard deviations.Tutor CoPilot was evaluated with approximately 900 tutors and 1,800 students in authentic tutoring sessions.A meta-analysis of 117 studies found an average critical-thinking intervention effect near 0.30 standard deviations.Four safer assessment safeguards are version histories, oral questioning, supervised baseline work, and source verification.
The value of the research is not only what each source knew, but what became visible when their evidence was combined.
En resumenMade for Eduardo Santiago, by HanademiFuentes: Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045.; Freeman, S., et al. (2014). Active learning increases studentperformance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences.; Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis.Educational Psychology Review, 30, 703-725.; Wang, R. E., et al. (2024). Tutor CoPilot: A human-AI approach for scaling real-time expertise. NBER Working Paper 33097.; Abrami, P. C., et al. (2008). Instructional…Un estudio bibliográfico encontró tasas de referencias inventadas del 55% para GPT-3.5 y del 18% para GPT-4.En 225 estudios universitarios STEM, el aprendizaje activo mejoró el rendimiento en exámenes en 0,47 desviaciones estándar.Un metaanálisis de 64 estudios estimó que la autoexplicación guiada mejoró el aprendizaje aproximadamente 0,55 desviaciones estándar.Tutor CoPilot fue evaluado con aproximadamente 900 tutores y 1.800 estudiantes en sesiones auténticas de tutoría.Un metaanálisis de 117 estudios encontró un efecto promedio cercano a 0,30 desviaciones estándar para intervenciones de pensamiento crítico.Cuatro salvaguardas de evaluación más seguras son historiales de versiones, preguntas orales, trabajo base supervisado y verificación de fuentes.
El valor de la investigación no está solo en cada fuente, sino en lo que apareció al combinar sus evidencias.

The research behind this deck

Fluent performance is not dependable competence. The core question is what remains when assistance disappears.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

Un rendimiento fluido no equivale a una competencia fiable. La pregunta central es qué permanece cuando desaparece la ayuda.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English

Related researchInvestigación relacionada