Hanademi

AI agents work best with boundariesLos agentes de IA funcionan mejor con límites · V7

32 slides · 24 min · 2026-08-04 language
ENES
theme
LightDark
view
DeckTableTalk
brand
HanademiPlatzi
AI agents work best withboundariesMade for Luis Fernando Navarrete Estrada, by Hanademi
  1. The measured speed gain is large, but the experiment covered one bounded task.
  2. AI can accelerate the wrong answer when users misjudge its capability.
  3. Agents add a decision loop around a model.
Los agentes de IA funcionanmejor con límitesMade for Luis Fernando Navarrete Estrada, by Hanademi
  1. La mejora medida es grande, pero el experimento cubrió una sola tarea delimitada.
  2. La IA puede acelerar una respuesta incorrecta cuando se juzga mal su capacidad.
  3. Los agentes añaden un ciclo de decisión alrededor de un modelo.
Copilot cut a coding task by 89.7 minutesAverage completion time in GitHub's randomized JavaScript experiment, minutes.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHubCopilot. GitHub.Unitminutes160.9Without Copilot89.7Minutes saved71.2With Copilot
Start with the benefit because it is real. Developers using Copilot finished the defined task in 71.2 minutes, compared with 160.9 minutes for controls. That saved 89.7 minutes. The rest of the deck asks what happens when the task becomes broader, riskier or less testable.
Copilot redujo 89,7 minutos una tarea deprogramaciónTiempo promedio de ejecución en el experimento aleatorizado de JavaScript de GitHub, minutos.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHubCopilot. GitHub.Unidadminutos160,9Sin Copilot89,7Minutos ahorrados71,2Con Copilot
Empecemos por el beneficio porque es real. Los desarrolladores con Copilot terminaron la tarea definida en 71,2 minutos, frente a 160,9 minutos del control. Ahorraron 89,7 minutos. El resto de la presentación pregunta qué ocurre cuando la tarea es más amplia, riesgosa o difícil de evaluar.
Three terms to knowMade for Luis Fernando Navarrete Estrada, by HanademiAI agentSoftware that observes, decides and acts toward a goal.ReActA method that alternates reasoning with actions and observations.Context windowThe amount of input a model can process at once.Prompt injectionMalicious instructions that try to redirect a model or agent.
A model produces an answer. An agent adds observation, decisions, tools and repeated evaluation. ReAct is one way to organize that loop. Context and prompt injection matter because more information and more permissions create both capability and risk.
Tres términos esencialesMade for Luis Fernando Navarrete Estrada, by HanademiAgente de IASoftware que observa, decide y actúa para alcanzar un objetivo.ReActUn método que alterna razonamiento con acciones y observaciones.Ventana de contextoLa cantidad de entrada que un modelo puede procesar a la vez.Inyección de promptsInstrucciones maliciosas que intentan desviar un modelo o agente.
Un modelo produce una respuesta. Un agente añade observación, decisiones, herramientas y evaluación repetida. ReAct es una forma de organizar ese ciclo. El contexto y la inyección de prompts importan porque más información y permisos crean capacidad y riesgo.
GAIA put humans at 92% and the assistantnear 15%Reported performance on GAIA tasks requiring reasoning, tools and access to information.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.UnitperformanceHumans92%GPT-4 assistant15%6.1x
An agent must do more than answer. It has to find information, use tools and keep track of a goal. GAIA reported about 92% for humans and about 15% for its strongest evaluated GPT-4-based assistant. That gap is the foundation for every design choice that follows.
GAIA situó a humanos en 92% y al asistentecerca de 15%Rendimiento informado en tareas GAIA que requieren razonamiento, herramientas y acceso a información.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.UnidadrendimientoHumanos92 %Asistente con GPT-415 %6,1x
Un agente debe hacer más que responder. Tiene que encontrar información, usar herramientas y seguir un objetivo. GAIA informó cerca de 92% para humanos y 15% para su asistente evaluado más fuerte basado en GPT-4. Esa brecha sustenta cada decisión de diseño que sigue.
Five ReAct cycles require at least 10 modelstagesMinimum reasoning and action stages implied by an alternating ReAct design.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Yao, S., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference onLearning Representations.Unitminimum model stages1 cycle23 cycles65 cycles10
ReAct behaves like navigating while checking signs after every turn. That feedback can help, but every loop adds work. One cycle needs at least 2 stages, while 5 cycles need at least 10. Agent depth should therefore be earned by the task, not added by default.
Cinco ciclos ReAct requieren al menos 10etapas del modeloEtapas mínimas de razonamiento y acción implicadas por un diseño ReAct alternante.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Yao, S., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference onLearning Representations.Unidadetapas mínimas del modelo1 ciclo23 ciclos65 ciclos10
ReAct se parece a navegar comprobando señales después de cada giro. Esa retroalimentación puede ayudar, pero cada ciclo añade trabajo. Un ciclo necesita al menos 2 etapas y 5 ciclos requieren al menos 10. La profundidad del agente debe justificarse por la tarea.
ChatGPT made writing 40% faster and 18%betterExperimental changes in task time and evaluated output quality among 453 professionals.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science,381, 187-192.Unitpercent changeLess task time40%Higher quality18%2.2x
This was not merely a speed effect. In a 453-person experiment, ChatGPT reduced writing time by 40% while evaluated quality increased by 18%. The result shows why bounded knowledge work is an attractive starting point. The task and evaluation method still define where the finding applies.
ChatGPT hizo la redacción 40% más rápida y18% mejorCambios experimentales en tiempo y calidad evaluada del resultado entre 453 profesionales.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence.Science, 381, 187-192.Unidadcambio porcentualMenos tiempo40 %Mayor calidad18 %2,2x
No fue solo un efecto de velocidad. En un experimento con 453 personas, ChatGPT redujo 40% el tiempo de redacción mientras la calidad evaluada aumentó 18%. El resultado explica por qué el trabajo intelectual delimitado es un buen punto de partida. La tarea y la evaluación definen dónde aplica.
GPT-4o raised central-bank work quality byup to 44%Randomized GPT-4o access at the National Bank of Slovakia, quality and completion-time improvements.Made for Luis Fernando Navarrete Estrada, by HanademiSources: cepr.org.Unitpercent improvementOutput quality44%Completion time21%
The productivity story extends beyond software and casual writing. Randomized GPT-4o access at the National Bank of Slovakia improved output quality by up to 44% and reduced completion time by 21%. The key is assistance inside a controlled institution, not unsupervised authority.
Sources
GPT-4o elevó hasta 44% la calidad deltrabajo en un banco centralAcceso aleatorio a GPT-4o en el Banco Nacional de Eslovaquia, mejoras de calidad y tiempo de ejecución.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: cepr.org.Unidadmejora porcentualCalidad del resultado44 %Tiempo de ejecución21 %
La historia de productividad va más allá del software y la redacción casual. El acceso aleatorio a GPT-4o en el Banco Nacional de Eslovaquia mejoró hasta 44% la calidad y redujo 21% el tiempo de ejecución. La clave es asistencia dentro de una institución controlada, no autoridad sin supervisión.
Fuentes
Novice support workers gained more thantwice the averageIncrease in resolved issues per hour across 5,179 support agents, average versus approximate novicegain.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work. National Bureau of Economic ResearchWorking Paper 31161.Unitproductivity increase34%Novices14%All agents
The average productivity gain was 14%, but novices gained about 34%. The most experienced workers saw minimal gains. This suggests that assistance can spread practices already held by stronger workers. It also means the value of a tool depends heavily on who uses it.
Los novatos en soporte mejoraron más deldoble que el promedioAumento de casos resueltos por hora entre 5.179 agentes de soporte, promedio frente a mejora aproximadade novatos.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work. National Bureau of Economic ResearchWorking Paper 31161.Unidadaumento de productividad34 %Novatos14 %Todos los agentes
La mejora promedio fue 14%, pero los novatos ganaron cerca de 34%. Los trabajadores más experimentados obtuvieron beneficios mínimos. Esto sugiere que la asistencia puede difundir prácticas que ya dominan los trabajadores fuertes. El valor depende mucho de quién usa la herramienta.
GPT-4 accelerated consultants but failedbeyond its frontierReported changes in task volume, completion speed and outside-frontier accuracy for consultants usingGPT-4.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.Task volume and speed are percent changes.Unitreported changeMore tasks12.2Faster work25.1Accuracy penalty19
Inside the technology frontier, consultants completed 12.2% more tasks and worked 25.1% faster. Outside it, they were 19 percentage points less likely to solve the task correctly. This is the central operating risk. Speed is valuable only when the system can detect which corridor it is in.
GPT-4 aceleró a consultores, pero fallófuera de su fronteraCambios informados en volumen, velocidad y precisión fuera de la frontera para consultores con GPT-4.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper24-013.El volumen y la velocidad son cambios porcentuales.Unidadcambio informadoMás tareas12,2Trabajo más rápido25,1Penalización de precisión19
Dentro de la frontera tecnológica, los consultores completaron 12,2% más tareas y trabajaron 25,1% más rápido. Fuera de ella, tuvieron 19 puntos porcentuales menos de probabilidad de responder correctamente. Este es el riesgo operativo central. La velocidad vale cuando el sistema detecta en qué corredor está.
Pakistan’s JudgeGPT experimentrandomized 1,559 judges across 118courts into three training arrangements.The experiment compared targeted training, generic training with AI and generic trainingwithout AI.Sources: elliottash.com.
Pakistan’s JudgeGPT experiment randomized 1,559 judges across 118 courts into three training arrangements.
El experimento JudgeGPT de Pakistánasignó aleatoriamente a 1.559 jueces de118 tribunales entre tres modalidadesde capacitación.El experimento comparó capacitación específica, capacitación genérica con IA ycapacitación genérica sin IA.Fuentes: elliottash.com.
El experimento JudgeGPT de Pakistán asignó aleatoriamente a 1.559 jueces de 118 tribunales entre tres modalidades de capacitación.
Researchers estimated that Pakistan’sJudgeGPT trial generated $38.50 in savingsfor every dollar spent.This reported estimate expresses savings relative to spending and comes from aseparate source.Sources: venturepost.co.
Researchers estimated that Pakistan’s JudgeGPT trial generated $38.50 in savings for every dollar spent.
Los investigadores estimaron que el ensayode JudgeGPT en Pakistán generó US $38,50en ahorros por cada dólar invertido.Esta estimación informada expresa los ahorros en relación con el gasto y proviene deuna fuente distinta.Fuentes: venturepost.co.
Los investigadores estimaron que el ensayo de JudgeGPT en Pakistán generó US$38,50 en ahorros por cada dólar invertido.
AI improved benefits caseworker accuracyby an average of 40%.The result concerns assistive use by caseworkers in a defined public-benefits setting.Sources: digitalgovernmenthub.org.
A randomized trial and a 14-week Los Angeles pilot tested a generative assistant for public-benefit caseworkers. Average accuracy improved by 40%. The tool helped workers navigate a difficult body of rules. It supported accountable staff rather than replacing their authority.
La IA mejoró en promedio 40% la precisiónde gestores de beneficios.El resultado se refiere al uso asistencial por gestores en un entorno definido debeneficios públicos.Fuentes: digitalgovernmenthub.org.
Un ensayo aleatorizado y un piloto de 14 semanas en Los Ángeles probaron un asistente generativo para gestores de beneficios. La precisión promedio mejoró 40%. La herramienta ayudó a navegar un conjunto complejo de reglas. Apoyó a personal responsable en vez de reemplazar su autoridad.
Copilot users reported faster repetitivework and better focusShare of GitHub survey respondents reporting 3 perceived benefits from Copilot.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Kalliamvakou, E. (2022). Research: Quantifying GitHub Copilot's impact on developer productivity and happiness. GitHub.Unitsurvey respondents96%Repetitive work faster88%More productive74%More satisfying work
The experience data support the experimental speed result. In GitHub's survey, 96% said repetitive work was faster, 88% felt more productive and 74% focused on more satisfying work. These are self-reported perceptions, but they show where assistance feels valuable to users.
Usuarios de Copilot informaron másvelocidad y mejor enfoqueProporción de participantes en una encuesta de GitHub que informó 3 beneficios percibidos de Copilot.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Kalliamvakou, E. (2022). Research: Quantifying GitHub Copilot's impact on developer productivity andhappiness. GitHub.Unidadparticipantes de la encuesta96 %Trabajo repetitivo más rápido88 %Más productivos74 %Trabajo más satisfactorio
Los datos de experiencia respaldan el resultado experimental de velocidad. En la encuesta de GitHub, 96% dijo que el trabajo repetitivo fue más rápido, 88% se sintió más productivo y 74% se concentró en trabajo más satisfactorio. Son percepciones autoinformadas, pero muestran dónde la asistencia aporta valor.
AI coding assistants saved an estimated 28working days per developer.The figure is an annualized estimate from a government trial, not a direct count of 28absent workdays.Sources: gov.uk.
A UK cross-government trial translated coding-assistant use into an annual estimate. Each participating developer saved the equivalent of 28 working days. That is enough to reshape delivery capacity if the saved time becomes useful output. The estimate still depends on how recorded time savings translate into a full year.
Sources
Los asistentes de programaciónahorraron 28 días laboralesestimados por desarrollador.La cifra es una estimación anualizada de un ensayo público, no un conteo directo de 28días de ausencia.Fuentes: gov.uk.
Un ensayo transversal del gobierno británico convirtió el uso de asistentes de programación en una estimación anual. Cada desarrollador participante ahorró el equivalente a 28 días laborales. Eso puede cambiar la capacidad de entrega si el tiempo liberado se convierte en resultados útiles. La cifra depende de cómo se anualicen los ahorros observados.
Fuentes
Codex reached 72.3% only after allowing100 attemptsReported Codex-12B pass rates on HumanEval as more generated samples were allowed; build-timeverification was incomplete.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Chen, M., et al. (2021). Evaluating large language models trained on code. OpenAI.UnitHumanEval pass ratepass@128.8%pass@1046.8%pass@10072.3%
One attempt produced a reported 28.8% pass rate. Allowing 10 attempts raised it to 46.8%, and 100 attempts raised it to 72.3%. The capability was real, but repeated generation and testing carried much of the result. Without a trustworthy evaluator, more drafts can simply multiply uncertainty.
Codex alcanzó 72,3% solo al permitir 100intentosTasas informadas de Codex-12B en HumanEval al permitir más muestras; la verificación durante laconstrucción fue incompleta.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Chen, M., et al. (2021). Evaluating large language models trained on code. OpenAI.Unidadtasa de aprobación en HumanEvalpass@128,8 %pass@1046,8 %pass@10072,3 %
Un intento produjo una tasa informada de 28,8%. Permitir 10 intentos la elevó a 46,8% y 100 intentos a 72,3%. La capacidad era real, pero la generación repetida y las pruebas sostuvieron gran parte del resultado. Sin un evaluador confiable, más borradores pueden multiplicar la incertidumbre.
Advertised context grew from thousands to2 million tokensSelected advertised context-window milestones for Claude, GPT and Gemini. Log scale: every gridline is×10.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Anthropic. (2023). Claude 2.1. Anthropic.; OpenAI. (2023). New models and developer products announced at DevDay. OpenAI.;Google. (2024). Gemini 1.5 Pro updates and two million token context. Google.The release stages align product milestones for comparison; they are not shared calendar dates.Unittokens1K10K100K1,000K10,000KEarlier releaseExpanded releaseLargest citedClaudeGPTGeminiGemini's advertised windowreached 2 million tokens.
Context windows grew by orders of magnitude across major model families. Claude reached 200,000 tokens, GPT reached 128,000 and Gemini reached 2 million in the cited releases. A larger desk can hold more documents. It still does not guarantee that the model will find the right page or act correctly on it.
El contexto anunciado creció de miles a 2millones de tokensHitos seleccionados de ventanas de contexto anunciadas para Claude, GPT y Gemini. Escala logarítmica:cada línea es ×10.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Anthropic. (2023). Claude 2.1. Anthropic.; OpenAI. (2023). New models and developer products announced at DevDay. OpenAI.; Google.(2024). Gemini 1.5 Pro updates and two million token context. Google.Las etapas alinean hitos de producto para compararlos; no representan fechas comunes.Unidadtokens1K10K100K1.000K10.000KVersión anteriorVersión ampliadaMayor valor citadoClaudeGPTGeminiLa ventana anunciada de Geminillegó a 2 millones de tokens.
Las ventanas de contexto crecieron por órdenes de magnitud en grandes familias de modelos. Claude llegó a 200.000 tokens, GPT a 128.000 y Gemini a 2 millones en las versiones citadas. Un escritorio mayor puede contener más documentos. No garantiza que el modelo encuentre la página correcta ni actúe bien.
Gemini's 90% benchmark edge was notapples-to-applesReported MMLU results for Gemini Ultra and GPT-4 used different prompting procedures.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Google DeepMind. (2023). Gemini: A family of highly capable multimodal models. Google DeepMind.; OpenAI. (2023). GPT-4technical report. OpenAI.Google reported the figures with different prompting procedures, weakening direct comparability.UnitMMLU score90%Gemini Ultra86.4%GPT-4
Gemini Ultra was reported at 90.0% and GPT-4 at 86.4% on MMLU. The comparison looks decisive until the evaluation procedures are examined. Different prompting methods weaken the direct comparison. More importantly, high exam scores do not establish reliable autonomous execution.
La ventaja de 90% de Gemini no fue unacomparación equivalenteLos resultados informados de MMLU para Gemini Ultra y GPT-4 usaron procedimientos de promptingdiferentes.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Google DeepMind. (2023). Gemini: A family of highly capable multimodal models. Google DeepMind.; OpenAI. (2023).GPT-4 technical report. OpenAI.Google informó las cifras con procedimientos de prompting diferentes, lo que debilita la comparación directa.Unidadpuntuación MMLU90 %Gemini Ultra86,4 %GPT-4
Gemini Ultra fue informado con 90,0% y GPT-4 con 86,4% en MMLU. La comparación parece decisiva hasta revisar los procedimientos. Métodos de prompting distintos debilitan la comparación directa. Además, puntuaciones académicas altas no establecen una ejecución autónoma confiable.
Llama grew to 405 billion parameters, witha license boundaryLargest released Llama model by generation, billions of parameters; Llama 3.1 products above 700 millionmonthly users require additional permission.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Meta AI. (2024). The Llama 3 herd of models. Meta AI.; Meta. (2024). Llama 3.1 Community License Agreement. Meta.Unitbillion parameters0200400Llama 1, 2023Llama 2, 2023Llama 3.1, 2024Llama 3.1 reached 405 billionparameters.65
Meta expanded its largest Llama release from 65 billion parameters to 405 billion. Downloadable weights can offer more deployment control than a closed service. That control comes with responsibilities for infrastructure, security and licensing. Llama 3.1 requires additional permission above 700 million monthly active users.
Llama creció a 405.000 millones deparámetros, con un límite de licenciaMayor modelo Llama por generación, miles de millones de parámetros; productos con más de 700 millonesde usuarios mensuales requieren permiso adicional.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Meta AI. (2024). The Llama 3 herd of models. Meta AI.; Meta. (2024). Llama 3.1 Community LicenseAgreement. Meta.Unidadmiles de millones de parámetros0200400Llama 1, 2023Llama 2, 2023Llama 3.1, 2024Llama 3.1 alcanzó 405.000millones de parámetros.65
Meta amplió su mayor lanzamiento de Llama de 65.000 millones a 405.000 millones de parámetros. Los pesos descargables pueden dar más control que un servicio cerrado. Ese control trae responsabilidades de infraestructura, seguridad y licencia. Llama 3.1 exige permiso adicional por encima de 700 millones de usuarios activos mensuales.
ChatGPT tripled weekly users in 13 monthsReported weekly active users, millions, November 2023 to December 2024.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Reuters. (2024). OpenAI's ChatGPT weekly users rise to 200 million. Reuters.; CNBC. (2024). OpenAI says ChatGPThas 300 million weekly active users. CNBC.Unitmillion weekly users0100200300Nov 2023Aug 2024Dec 2024Reported use reached 300million weekly users.
Reported weekly use rose from 100 million to 300 million in 13 months. A complex technology reached mass audiences through a simple conversation interface. That scale multiplies both useful assistance and confidently wrong output. Verification cannot remain an expert-only practice.
ChatGPT triplicó sus usuarios semanales en13 mesesUsuarios activos semanales informados, millones, de noviembre de 2023 a diciembre de 2024.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Reuters. (2024). OpenAI's ChatGPT weekly users rise to 200 million. Reuters.; CNBC. (2024). OpenAIsays ChatGPT has 300 million weekly active users. CNBC.Unidadmillones de usuarios semanales0100200300Nov 2023Ago 2024Dic 2024El uso informado llegó a 300millones de usuariossemanales.
El uso semanal informado creció de 100 millones a 300 millones en 13 meses. Una tecnología compleja llegó a audiencias masivas mediante una interfaz conversacional sencilla. Esa escala multiplica tanto la asistencia útil como las respuestas incorrectas y seguras. Verificar ya no es una tarea exclusiva de expertos.
MD Anderson spent about $62.1 millionbefore pausing its oncology project.The figure concerns MD Anderson's Oncology Expert Advisor project and should not begeneralized to every IBM Watson deployment.Sources: University of Texas System Audit Office. (2016). Special review of procurement procedures related to the Oncology Expert Advisor project.University of Texas System.
Laboratory capability is only one part of deployment. A University of Texas review reported about $62.1 million spent on MD Anderson's Oncology Expert Advisor before the project was placed on hold. Integration, procurement and operating rules can dominate the outcome. A powerful engine cannot compensate for incompatible roads.
MD Anderson gastó cerca deUS$62,1 millones antes desuspender su proyecto oncológico.La cifra corresponde al proyecto Oncology Expert Advisor de MD Anderson y no debegeneralizarse a todos los despliegues de IBM Watson.Fuentes: University of Texas System Audit Office. (2016). Special review of procurement procedures related to the Oncology Expert Advisor project.University of Texas System.
La capacidad de laboratorio es solo una parte del despliegue. Una revisión de la Universidad de Texas informó cerca de US$62,1 millones gastados en Oncology Expert Advisor antes de suspender el proyecto. La integración, la contratación y las reglas operativas pueden dominar el resultado. Un motor potente no compensa carreteras incompatibles.
OSWorld exposed weak agent performanceon real computer tasksReported OSWorld success rates for humans and the strongest evaluated GPT-4V-based agent; liveverification was incomplete.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Xie, T., et al. (2024). OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. NeuralInformation Processing Systems.Unittask successHumans72.4%GPT-4V agent12.2%
The reported gap is not a small error rate. Humans achieved 72.4% success, while the strongest evaluated GPT-4V-based agent achieved 12.2%. That makes broad permissions hard to justify. Capability limits and permission limits should be designed together.
OSWorld reveló un bajo rendimiento deagentes en tareas informáticas realesTasas informadas de éxito en OSWorld para humanos y el agente evaluado más fuerte con GPT-4V; laverificación en vivo fue incompleta.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Xie, T., et al. (2024). OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. NeuralInformation Processing Systems.Unidadéxito en tareasHumanos72,4 %Agente con GPT-4V12,2 %
La brecha informada no es una tasa pequeña de error. Los humanos lograron 72,4% de éxito y el agente evaluado más fuerte con GPT-4V obtuvo 12,2%. Eso dificulta justificar permisos amplios. Los límites de capacidad y permisos deben diseñarse juntos.
OWASP ranked prompt injection as the topLLM riskTop 3 risks in the OWASP Top 10 for Large Language Model Applications 2023; lower rank is more severe.Made for Luis Fernando Navarrete Estrada, by HanademiSources: OWASP. (2023). OWASP Top 10 for Large Language Model Applications 2023. OWASP.Unitrisk rankPrompt injection1Insecure output handling2Training-data poisoning3
Prompt injection ranked first in OWASP's 2023 list, ahead of insecure output handling and training-data poisoning. A chatbot can produce a bad sentence. An agent with tools can turn a bad instruction into an action. Least privilege, logging and approval gates are therefore product features, not paperwork.
OWASP clasificó la inyección de promptscomo el principal riesgoLos 3 principales riesgos de OWASP para aplicaciones con modelos de lenguaje en 2023; un rango menorindica mayor severidad.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: OWASP. (2023). OWASP Top 10 for Large Language Model Applications 2023. OWASP.Unidadposición de riesgoInyección de prompts1Manejo inseguro de salidas2Envenenamiento de datos3
La inyección de prompts ocupó el primer lugar en la lista de OWASP de 2023, por delante del manejo inseguro de salidas y el envenenamiento de datos. Un chatbot puede producir una frase incorrecta. Un agente con herramientas puede convertir una instrucción incorrecta en una acción. El privilegio mínimo, los registros y las aprobaciones son funciones del producto.
A small multiagent debate creates at least 9responsesMinimum agent responses for one single-pass answer versus 3 agents debating for 3 rounds; build-timeverification was incomplete.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Du, Y., et al. (2024). Improving factuality and reasoning in language models through multiagent debate. InternationalConference on Machine Learning.Unitagent responsesSingle pass13 agents, 3 rounds9
Three agents debating for three rounds produce at least 9 responses instead of one. That can create specialization or disagreement. It also multiplies computation and coordination. If every agent shares the same model and blind spots, the extra voices may not provide independent verification.
Un pequeño debate multiagente crea almenos 9 respuestasRespuestas mínimas para una sola pasada frente a 3 agentes durante 3 rondas; la verificación durante laconstrucción fue incompleta.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Du, Y., et al. (2024). Improving factuality and reasoning in language models through multiagent debate. InternationalConference on Machine Learning.Unidadrespuestas de agentesUna sola pasada13 agentes, 3 rondas9
Tres agentes debatiendo durante tres rondas producen al menos 9 respuestas en vez de una. Eso puede crear especialización o desacuerdo. También multiplica el cómputo y la coordinación. Si todos comparten el mismo modelo y puntos ciegos, las voces adicionales pueden no aportar verificación independiente.
A three-agent debate uses nine times themodel work of one passTreat each agent response or model stage as one minimum model work unit and compare single pass, ReActcycles and a three-agent, three-round debate.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Yao, S., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference onLearning Representations.; Synergizing Reasoning and Acting in Language Models.; Du, Y., et al. (2024). Improvingfactuality and reasoning in language models through multiagent debate. International Conference on Machine Learning.UnitMinimum model generations or stagesSingle pass1ReAct, 1 cycle2ReAct, 3 cycles6Debate, 3 agents × 3 rounds9ReAct, 5 cycles10
The nine-response debate nearly matches the ten minimum stages required by five ReAct cycles, before counting coordination or synthesis.
Un debate de tres agentes usa nueve vecesel trabajo de una sola pasadaTratar cada respuesta de agente o etapa del modelo como una unidad mínima de trabajo y comparar unasola pasada, ciclos ReAct y un debate de tres agentes durante tres rondas.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Yao, S., et al. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference onLearning Representations.; Synergizing Reasoning and Acting in Language Models.; Du, Y., et al. (2024). Improvingfactuality and reasoning in language models through multiagent debate. International Conference on Machine Learning.UnidadGeneraciones o etapas mínimas del modeloUna sola pasada1ReAct, 1 ciclo2ReAct, 3 ciclos6Debate, 3 agentes × 3 rondas9ReAct, 5 ciclos10
El debate de nueve respuestas casi iguala las diez etapas mínimas requeridas por cinco ciclos ReAct, antes de contar la coordinación o la síntesis.
Use AI where tasks have clear inputs, testsand owners, then limit permissions andrequire approval for consequential actions.The evidence combines task fit, measured outcomes, advertised capability, execution riskand security priorities. It supports disciplined adoption, not treating fluency or scale asproof of reliability.Sources: Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. GitHub.; The Impact of AI on Productivity - Marginal REVOLUTION.;Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.; Navigating the Jagged Technological Frontier: Field | ResearchBunny.; Google. (2024). Ournext-generation model: Gemini 1.5. Google.; Google. (2024). Gemini 1.5 Pro updates and two million token context. Google.; Gemini 1.5 Pro 2M context window, code execution capabilities, and Gemma 2 are available…
The evidence supports using AI, but not treating fluency, scale or a single productivity result as reliability. Start with bounded tasks, test outcomes, restrict permissions and expand autonomy only when measured performance justifies it.
Use IA cuando las tareas tengan entradas,pruebas y responsables claros; limite lospermisos y exija aprobación para accionesimportantes.La evidencia combina adecuación de tarea, resultados medidos, capacidad anunciada,riesgo de ejecución y prioridades de seguridad. Respalda una adopción disciplinada, notratar la fluidez ni la escala como prueba de confiabilidad.Fuentes: Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. GitHub.; The Impact of AI on Productivity - Marginal REVOLUTION.;Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.; Navigating the Jagged Technological Frontier: Field | ResearchBunny.; Google. (2024). Ournext-generation model: Gemini 1.5. Google.; Google. (2024). Gemini 1.5 Pro updates and two million token context. Google.; Gemini 1.5 Pro 2M context window, code execution capabilities, and Gemma 2 are available…
La evidencia respalda el uso de IA, pero no confundir fluidez, escala ni un solo resultado de productividad con confiabilidad. Empiece con tareas delimitadas, evalúe resultados, restrinja permisos y amplíe la autonomía solo cuando el desempeño medido lo justifique.
GPT and Gemini expanded advertised context62.5-foldDivide each later advertised context window by the earliest stated window for the same model family.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Anthropic. (2023). Introducing 100K context windows. Anthropic.; Anthropic. (2023). Claude 2.1.Anthropic.; Introducing 100K Context Windows.UnitMultiple of earliest stated context windowClaude22.2GPT62.5Gemini62.5
GPT grew from 2,048 to 128,000 tokens and Gemini from 32,000 to two million, both 62.5-fold. Claude's move from about 9,000 to 200,000 tokens equals 22.2-fold.
GPT y Gemini ampliaron 62,5 veces elcontexto anunciadoDividir cada ventana de contexto anunciada posteriormente entre la primera ventana indicada para la mismafamilia de modelos.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Anthropic. (2023). Introducing 100K context windows. Anthropic.; Anthropic. (2023). Claude 2.1.Anthropic.; Introducing 100K Context Windows.UnidadMúltiplo de la primera ventana de contexto indicadaClaude22,2GPT62,5Gemini62,5
GPT pasó de 2.048 a 128.000 tokens y Gemini de 32.000 a dos millones, ambos 62,5 veces. El avance de Claude desde cerca de 9.000 hasta 200.000 tokens equivale a 22,2 veces.
Speed gains can coexist with either betteror worse outputMade for Luis Fernando Navarrete Estrada, by HanademiSources: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificialintelligence. Science, 381, 187-192.; Experimental evidence on the productivity effects of generative ....; Dell'Acqua, F.,et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.UnitReported changes, percent or percentage pointsOutput quality or correctness change, pointsNo output change40200-2001020304050Reported time or speed gain, percentSlovak central bank21% faster, up to 44% betterWriting study40% less time, 18% betterConsultants outside the frontier25.1% faster19 points less accurate
Within a 21% to 40% speed-gain range, measured outcomes span from a 44% quality increase to a 19-point correctness penalty outside the task frontier.
Las mejoras de velocidad pueden coexistircon resultados mejores o peoresMade for Luis Fernando Navarrete Estrada, by HanademiFuentes: Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificialintelligence. Science, 381, 187-192.; Experimental evidence on the productivity effects of generative ....; Dell'Acqua, F.,et al. (2023). Navigating the jagged technological frontier. Harvard Business School Working Paper 24-013.UnidadCambios informados, porcentaje o puntos porcentualesCambio en calidad o corrección, puntosSin cambio en el resultado40200-2001020304050Mejora informada de tiempo o velocidad, porcentajeBanco central de Eslovaquia21% más rápido, hasta 44% mejorEstudio de redacción40% menos tiempo, 18% mejorConsultores fuera de la frontera25,1% más rápidos19 puntos menos precisos
Dentro de un intervalo de mejora de velocidad del 21% al 40%, los resultados medidos van desde un aumento de calidad del 44% hasta una penalización de 19 puntos en corrección fuera de la frontera de tareas.
A 40% vulnerability rate halves Copilot'simplied speed advantageTreat 60% as the secure-output probability and divide Copilot's 71.2 minutes by 0.60, producing 118.7expected minutes per accepted program.Made for Luis Fernando Navarrete Estrada, by HanademiSources: Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidencefrom GitHub Copilot. GitHub.; The Impact of AI on Productivity - Marginal REVOLUTION.; Pearce, H., et al. (2022). Asleep atthe keyboard? Assessing the security of GitHub Copilot's code contributions. IEEE Symposium on Security and Privacy.UnitExpected minutes per accepted programControl baseline160.9Copilot, security-adjusted118.7Copilot, observed time71.2
The observed time advantage falls from 55.8% to 26.2% under this stress test. It assumes independent reruns, a secure control output and transfer of the security-study rate.
Una vulnerabilidad del 40% reduce a lamitad la ventaja implícita de CopilotTratar el 60% como la probabilidad de una salida segura y dividir los 71,2 minutos de Copilot entre 0,60, loque produce 118,7 minutos esperados por programa aceptado.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidencefrom GitHub Copilot. GitHub.; The Impact of AI on Productivity - Marginal REVOLUTION.; Pearce, H., et al. (2022). Asleep atthe keyboard? Assessing the security of GitHub Copilot's code contributions. IEEE Symposium on Security and Privacy.UnidadMinutos esperados por programa aceptadoBase de control160,9Copilot, ajustado por seguridad118,7Copilot, tiempo observado71,2
La ventaja de tiempo observada cae del 55,8% al 26,2% bajo esta prueba de tensión. Supone repeticiones independientes, una salida de control segura y la transferencia de la tasa del estudio de seguridad.
High knowledge scores do not translate intoautonomous task completionMade for Luis Fernando Navarrete Estrada, by HanademiSources: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; GAIA: a benchmark for general AI assistants | Research.; Zhou, S., et al. (2024). WebArena: Arealistic web environment for building autonomous agents. International Conference on Learning Representations.UnitReported benchmark score, percentGPT-4 MMLU86.4%17.4% of MMLU16.7% of MMLU14.1% of MMLU15.0GAIA task success15.0%14.4WebArena task success14.4%12.2OSWorld task success12.2%
Agent task scores equal only 14.1% to 17.4% of the MMLU reference. The ratios are a transfer index, not a direct conversion between benchmarks.
Las puntuaciones altas de conocimiento nose traducen en ejecución autónomaMade for Luis Fernando Navarrete Estrada, by HanademiFuentes: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; GAIA: a benchmark for general AI assistants | Research.; Zhou, S., et al. (2024). WebArena: Arealistic web environment for building autonomous agents. International Conference on Learning Representations.UnidadPuntuación informada en la evaluación, porcentajeMMLU de GPT-486,4%17,4% de MMLU16,7% de MMLU14,1% de MMLU15,0Éxito en tareas de GAIA15,0%14,4Éxito en tareas de WebArena14,4%12,2Éxito en tareas de OSWorld12,2%
Las puntuaciones de tareas de los agentes equivalen solo al 14,1% o 17,4% de la referencia MMLU. Las proporciones son un índice de transferencia, no una conversión directa entre evaluaciones.
Leading agents achieve about one-sixth ofhuman task successMade for Luis Fernando Navarrete Estrada, by HanademiSources: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; GAIA: a benchmark for general AI assistants | Research.; Zhou, S., et al. (2024). WebArena: Arealistic web environment for building autonomous agents. International Conference on Learning Representations.UnitTask success rate, percentTask success rate, percent0255075100GAIAAgent 15.0Human 92.016.3% of humanWebArenaAgent 14.4Human 78.218.4% of humanOSWorldAgent 12.2Human 72.416.9% of human
The agent-to-human ratios are 16.3%, 18.4% and 16.9%, a narrow range across three distinct task environments.
Los agentes líderes logran cerca de unasexta parte del éxito humanoMade for Luis Fernando Navarrete Estrada, by HanademiFuentes: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; GAIA: a benchmark for general AI assistants | Research.; Zhou, S., et al. (2024). WebArena: Arealistic web environment for building autonomous agents. International Conference on Learning Representations.UnidadTasa de éxito en tareas, porcentajeTasa de éxito en tareas, porcentaje0255075100GAIAAgente 15,0Humano 92,016,3% del humanoWebArenaAgente 14,4Humano 78,218,4% del humanoOSWorldAgente 12,2Humano 72,416,9% del humano
Las proporciones entre agente y humano son 16,3%, 18,4% y 16,9%, un intervalo estrecho en tres entornos de tareas distintos.
Prompt design covers only 6 of 12 specifiedcontrol elementsMade for Luis Fernando Navarrete Estrada, by HanademiSources: OpenAI. (2024). Prompt engineering best practices for ChatGPT. OpenAI.; Anthropic. (2024). Promptengineering overview. Anthropic.; NIST. (2023). Artificial Intelligence Risk Management Framework 1.0. U.S.Department of Commerce.UnitSpecified design elementsPrompt specification6 of 12 | 50%System safeguards3 of 12 | 25%Permission architecture3 of 12 | 25%
Prompt structure accounts for 50% of the combined checklist. Retrieval, protected execution, evaluation and permission architecture make up the other half.
El diseño del prompt cubre solo 6 de 12elementos de controlMade for Luis Fernando Navarrete Estrada, by HanademiFuentes: OpenAI. (2024). Prompt engineering best practices for ChatGPT. OpenAI.; Anthropic. (2024). Promptengineering overview. Anthropic.; NIST. (2023). Artificial Intelligence Risk Management Framework 1.0. U.S.Department of Commerce.UnidadElementos de diseño especificadosEspecificación del prompt6 de 12 | 50%Salvaguardas del sistema3 de 12 | 25%Arquitectura de permisos3 de 12 | 25%
La estructura del prompt representa el 50% de la lista combinada. La recuperación, la ejecución protegida, la evaluación y la arquitectura de permisos forman la otra mitad.
GOV.UK usage implies 390,000 questionsacross 150,000 officersCalculate 2.6 questions per user from 26,000 questions and 10,000 users, then apply that rate topopulations of 50,000 and 150,000.Made for Luis Fernando Navarrete Estrada, by HanademiSources: insidegovuk.blog.gov.uk.; agenticwire.news.UnitQuestions over 18 months0200K400K10,000 users50,000 users150,000 officers26,000390,000
At the observed rate, the Singapore officer base would generate 390,000 questions over 18 months, or about 21,700 per month. This is a scale counterfactual across different public services.
El uso de GOV.UK implica 390.000preguntas entre 150.000 funcionariosCalcular 2,6 preguntas por usuario a partir de 26.000 preguntas y 10.000 usuarios, y aplicar esa tasa apoblaciones de 50.000 y 150.000 personas.Made for Luis Fernando Navarrete Estrada, by HanademiFuentes: insidegovuk.blog.gov.uk.; agenticwire.news.UnidadPreguntas durante 18 meses0200K400K10.000 usuarios50.000 usuarios150.000 funcionarios26.000390.000
Con la tasa observada, la base de funcionarios de Singapur generaría 390.000 preguntas en 18 meses, cerca de 21.700 al mes. Es un contrafactual de escala entre servicios públicos distintos.
In summaryMade for Luis Fernando Navarrete Estrada, by HanademiSources: venturepost.co.; Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; elliottash.com.; Meta. (2024). Llama 3.1 Community License Agreement. Meta.; insidegovuk.blog.gov.uk.; OpenAI. (2023). GPT-4technical report. OpenAI.Researchers estimated that Pakistan’s JudgeGPT trial generated $38.50 in savings for every dollar spent.GAIA reported approximately 92% human performance, versus about 15% for its strongest evaluated GPT-4-based assistant.Pakistan’s JudgeGPT experiment randomized 1,559 judges across 118 courts into targeted-training,generic-training-with-AI and generic-training-without-AI arms.Llama 3.1's community license requires additional permission for products exceeding 700 million monthly active users.Across two GOV.UK Chat pilots, more than 10,000 users asked 26,000 government-service questions over 18 months.OpenAI reported GPT-4 at 86.4% on five-shot MMLU while also documenting hallucinations and confidently incorrect predictions.
The value of the research is not only what each source knew, but what became visible when their evidence was combined.
En resumenMade for Luis Fernando Navarrete Estrada, by HanademiFuentes: venturepost.co.; Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; elliottash.com.; Meta. (2024). Llama 3.1 Community License Agreement. Meta.; insidegovuk.blog.gov.uk.; OpenAI. (2023). GPT-4technical report. OpenAI.Los investigadores estimaron que el ensayo de JudgeGPT en Pakistán generó US$38,50 en ahorros por cada dólar invertido.GAIA informó aproximadamente 92% de rendimiento humano, frente a cerca de 15% para su asistente evaluado más fuerte basado en GPT-4.El experimento JudgeGPT de Pakistán asignó aleatoriamente a 1.559 jueces de 118 tribunales a tres grupos con distintascombinaciones de IA y capacitación.La licencia comunitaria de Llama 3.1 exige permiso adicional para productos que superen 700 millones de usuarios activos mensuales.En dos pilotos de GOV.UK Chat, más de 10.000 usuarios formularon 26.000 preguntas sobre servicios públicos durante 18 meses.OpenAI informó 86,4% para GPT-4 en MMLU con cinco ejemplos, mientras también documentó alucinaciones ypredicciones incorrectas expresadas con confianza.
El valor de la investigación no está solo en cada fuente, sino en lo que apareció al combinar sus evidencias.

The research behind this deck

The measured speed gain is large, but the experiment covered one bounded task. Agents add a decision loop around a model.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

La mejora medida es grande, pero el experimento cubrió una sola tarea delimitada. Los agentes añaden un ciclo de decisión alrededor de un modelo.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English

Related researchInvestigación relacionada