Hanademi

AI helps until it doesn'tLa IA ayuda hasta que deja de hacerlo · V7

25 slides · 13 min · 2026-08-11 language
ENES
theme
LightDark
view
TalkTable
brand
HanademiPlatzi
AI helps until it doesn'tMade for Raúl Mata Meneses, by Hanademi
  1. AI fluency does not prove understanding, truth, or sound judgment.
  2. Within compatible tasks, GPT-4 raised average quality by roughly 40%.
  3. The best tool is the one that passes an internal test with acceptable risks and costs.
La IA ayuda hasta que deja dehacerloMade for Raúl Mata Meneses, by Hanademi
  1. La fluidez de una IA no demuestra comprensión, verdad ni buen juicio.
  2. Dentro de tareas compatibles, GPT-4 elevó aproximadamente 40% la calidad promedio.
  3. La mejor herramienta es la que supera una prueba interna con riesgos y costos aceptables.
Within its boundary, GPT-4 raised quality40%Changes observed in 758 consultants who used GPT-4 on tasks within the evaluated boundary, 2023.Made for Raúl Mata Meneses, by HanademiSources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AIon knowledge worker productivity and quality. Harvard Business School.Unitpercentage change40%Quality25.1%Speed12.2%Tasks completed
This experiment tracked 758 consultants on real knowledge work tasks. Within the evaluated boundary, they completed more work, moved faster, and raised quality by roughly 40%. The opportunity is large, but depends on recognizing where the tool works.
Dentro de su frontera, GPT-4 elevó lacalidad 40%Cambios observados en 758 consultores que usaron GPT-4 en tareas dentro de la frontera evaluada, 2023.Made for Raúl Mata Meneses, by HanademiFuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AIon knowledge worker productivity and quality. Harvard Business School.Unidadcambio porcentual40 %Calidad25,1 %Velocidad12,2 %Tareas completadas
Este experimento siguió a 758 consultores en tareas reales de conocimiento. Dentro de la frontera evaluada, hicieron más trabajo, avanzaron más rápido y elevaron cerca de 40% la calidad. La oportunidad es grande, pero depende de reconocer dónde funciona la herramienta.
Five terms to speak the same languageMade for Raúl Mata Meneses, by HanademiModelMathematical engine that recognizes patterns and generates results.APIConnection that lets you integrate a model into another application.TokenSmall unit of text that the model processes and bills.BenchmarkStandardized test to compare capabilities under defined conditions.HallucinationFalse or wrong response presented with apparent confidence.
These words appear throughout the module. A model is the engine, while a tool adds interface, data, permissions, and integrations. Benchmarks help compare, but never represent all real work.
Cinco términos para hablar el mismoidiomaMade for Raúl Mata Meneses, by HanademiModeloMotor matemático que reconoce patrones y genera resultados.APIConexión que permite integrar un modelo en otra aplicación.TokenUnidad pequeña de texto que el modelo procesa y cobra.BenchmarkPrueba estandarizada para comparar capacidades bajo condiciones definidas.ConfabulaciónRespuesta falsa o errónea presentada con seguridad aparente.
Estas palabras aparecen durante todo el módulo. Un modelo es el motor, mientras una herramienta agrega interfaz, datos, permisos e integraciones. Los benchmarks ayudan a comparar, pero nunca representan todo el trabajo real.
HumanEval rose from 26.2% to 84.1% in twoyearsBest reported result on a code generation test, 2021 and 2023.Made for Raúl Mata Meneses, by HanademiSources: Stanford Institute for Human-Centered Artificial Intelligence. (2024). Artificial Intelligence Index Report 2024.; NationalInstitute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial IntelligenceProfile.Unitpercentage202126.2%202384.1%3.2x
The jump in HumanEval shows how much models improved at generating code in short time. Still, this exam measures a narrow capability. A fluent response can still be false, incomplete, or wrong for the situation.
HumanEval pasó de 26,2% a 84,1% en dosañosMejor resultado reportado en una prueba de generación de código, 2021 y 2023.Made for Raúl Mata Meneses, by HanademiFuentes: Stanford Institute for Human-Centered Artificial Intelligence. (2024). Artificial Intelligence Index Report 2024.; National Instituteof Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile.Unidadporcentaje202126,2 %202384,1 %3,2x
El salto en HumanEval muestra cuánto mejoraron los modelos para generar código en poco tiempo. Aun así, este examen mide una capacidad estrecha. Una respuesta fluida puede seguir siendo falsa, incompleta o inadecuada para la situación.
With AI, consultants were 19 percentagepoints less likely to get it right.The result corresponds to a specific task in the experiment and should not be generalizedto all tasks or tools.Sources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge workerproductivity and quality. Harvard Business School.
The same tool that accelerated work produced the opposite effect on a task outside its boundary. Participants relied on a capability that did not fit the problem. The risk is not only that AI fails, but that it advances a wrong answer faster.
Con IA, los consultores fueron 19 puntosporcentuales menos propensos a acertar.El resultado corresponde a una tarea específica del experimento y no debe generalizarsea todas las tareas ni herramientas.Fuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge workerproductivity and quality. Harvard Business School.
La misma herramienta que aceleró el trabajo produjo el efecto contrario en una tarea fuera de su frontera. Los participantes confiaron en una capacidad que no encajaba con el problema. El riesgo no es solo que la IA falle, sino que haga avanzar más rápido una respuesta equivocada.
The engine became cheaper, the tool keptcharging for accessPublished prices for API input per million tokens and monthly subscriptions per user, 2023 and 2024.Made for Raúl Mata Meneses, by HanademiSources: OpenAI. (2023). GPT-4 API general availability and deprecation of older models in the Completions API.; OpenAI.(2023). Introducing ChatGPT Plus.; Microsoft. (2023). Introducing Microsoft 365 Copilot pricing for commercial customers.UnitUSD, separate scalesEngine price$30$10$5GPT-4GPT-4 TurboGPT-4oTool price$20$30ChatGPT PlusMicrosoft 365 Copilot
The cost of GPT engine input fell from USD 30 to USD 5 per million tokens. That did not make the price of applications presenting it to the user disappear. A tool can maintain its brand, change the internal model, and charge for permissions, continuity, and integration.
El motor se abarató, la herramienta siguiócobrando accesoPrecios publicados para entrada de API por millón de tokens y suscripciones mensuales por usuario, 2023y 2024.Made for Raúl Mata Meneses, by HanademiFuentes: OpenAI. (2023). GPT-4 API general availability and deprecation of older models in the Completions API.; OpenAI.(2023). Introducing ChatGPT Plus.; Microsoft. (2023). Introducing Microsoft 365 Copilot pricing for commercialcustomers.UnidadUSD, escalas separadasPrecio del motor$30$10$5GPT-4GPT-4 TurboGPT-4oPrecio de la herramienta$20$30ChatGPT PlusMicrosoft 365 Copilot
El costo de entrada del motor GPT cayó de USD 30 a USD 5 por millón de tokens. Eso no hizo desaparecer el precio de las aplicaciones que lo presentan al usuario. Una herramienta puede mantener su marca, cambiar el modelo interno y cobrar por permisos, continuidad e integración.
The ecosystem crosses brandsTwo speakers on a stage with the OpenAI and Microsoft logos behind them.Sources: Copilot Pro vs. ChatGPT Plus: Microsoft's new paid service offers alternative to OpenAI subscription – GeekWire. geekwire.com.
OpenAI and Microsoft sharing a stage reinforces why model providers, product brands, and workplace tools must be distinguished.
El ecosistema cruza marcasDos ponentes en un escenario con los logotipos de OpenAI y Microsoft al fondo.Fuentes: Copilot Pro vs. ChatGPT Plus: Microsoft's new paid service offers alternative to OpenAI subscription – GeekWire. geekwire.com.
OpenAI y Microsoft compartiendo un escenario refuerzan por qué debemos distinguir proveedores de modelos, marcas de productos y herramientas de trabajo.
Copilot's theoretical time saving breakseven at $48 per labor hourMade for Raúl Mata Meneses, by HanademiSources: Microsoft. (2023). Introducing Microsoft 365 Copilot pricing for commercial customers.; Microsoft SetsExpensive Price Tag for New Corporate AI Products - Bloomberg.; Microsoft. (2023). What can Copilot's earliestusers teach us about generative AI at work?Unitannual net value per user, USDNet value after the $360 annual licenseZero net value-$173 per year$25 labor/hr+$15 per year$50 labor/hr+$390 per year$100 labor/hrBreak-even: $48 per hour
Under the reported saving, a $25 labor hour loses about $173 annually, a $50 hour barely clears the cost, and a $100 hour produces $390 before implementation and review costs.
El ahorro teórico de Copilot alcanza elequilibrio con una hora laboral de US$48Made for Raúl Mata Meneses, by HanademiFuentes: Microsoft. (2023). Introducing Microsoft 365 Copilot pricing for commercial customers.; Microsoft SetsExpensive Price Tag for New Corporate AI Products - Bloomberg.; Microsoft. (2023). What can Copilot's earliestusers teach us about generative AI at work?Unidadvalor neto anual por usuario, USDValor neto después de la licencia anual de US$360Valor neto cero-US$173 al añoUS$25/h laboral+US$15 al añoUS$50/h laboral+US$390 al añoUS$100/h laboralEquilibrio: US$48 por hora
Con el ahorro reportado, una hora laboral de US$25 pierde cerca de US$173 al año, una de US$50 apenas supera el costo y una de US$100 produce US$390 antes de costos de implantación y revisión.
ChatGPT added 100 million weekly users infour monthsWeekly users reported by OpenAI, August and December 2024, in millions.Made for Raúl Mata Meneses, by HanademiSources: Reuters. (2024). OpenAI says ChatGPT's weekly users have grown to 200 million.; OpenAI. (2024). Twelve daysof OpenAI.; OpenAI. (2024). GPT-4o and more tools to ChatGPT free.Unitmillions of weekly usersAugust 2024200December 20243001.5x
ChatGPT grew from 200 to 300 million weekly users reported in just four months. That scale enables learning, support, and recognition within teams. However, the free plan maintains variable limits and does not guarantee continuous capacity for intensive work.
ChatGPT sumó 100 millones de usuariossemanales en cuatro mesesUsuarios semanales reportados por OpenAI, agosto y diciembre de 2024, en millones.Made for Raúl Mata Meneses, by HanademiFuentes: Reuters. (2024). OpenAI says ChatGPT's weekly users have grown to 200 million.; OpenAI. (2024).Twelve days of OpenAI.; OpenAI. (2024). GPT-4o and more tools to ChatGPT free.Unidadmillones de usuarios semanalesAgosto de 2024200Diciembre de 20243001,5x
ChatGPT creció de 200 a 300 millones de usuarios semanales reportados en solo cuatro meses. Esa escala facilita aprendizaje, soporte y reconocimiento dentro de los equipos. Sin embargo, el plan gratuito mantiene límites variables y no garantiza capacidad continua para trabajo intensivo.
Access comes through an interfaceA phone displays the OpenAI logo against a background of generative AI text.Sources: ChatGPT added 50 million weekly users in just two months. engadget.com.
A familiar phone screen helps turn an abstract AI category into a recognizable product touchpoint.
El acceso llega mediante una interfazUn teléfono muestra el logotipo de OpenAI sobre un fondo con texto acerca de la IA generativa.Fuentes: ChatGPT added 50 million weekly users in just two months. engadget.com.
La pantalla familiar de un teléfono ayuda a convertir una categoría abstracta de IA en un punto de contacto reconocible.
GPT-4o lowered voice response from 2.8seconds to 320 msAudio latency reported by OpenAI for GPT-4o versus the previous voice system, 2024.Made for Raúl Mata Meneses, by HanademiSources: OpenAI. (2024). Hello GPT-4o.UnitmillisecondsGPT-4o minimum232GPT-4o average320Previous system2,800
A tool's experience depends on more than text quality. GPT-4o reduced the reported average voice response time from 2.8 seconds to 320 milliseconds. That difference changes how it feels to use the product, though it does not demonstrate greater accuracy.
GPT-4o bajó la respuesta de voz de 2,8segundos a 320 msLatencia de audio reportada por OpenAI para GPT-4o frente al sistema de voz anterior, 2024.Made for Raúl Mata Meneses, by HanademiFuentes: OpenAI. (2024). Hello GPT-4o.UnidadmilisegundosMínimo GPT-4o232Promedio GPT-4o320Sistema anterior2.800
La experiencia de una herramienta depende de más que la calidad del texto. GPT-4o redujo el promedio reportado de respuesta de voz desde 2,8 segundos hasta 320 milisegundos. Esa diferencia cambia cómo se siente usar el producto, aunque no demuestra mayor exactitud.
Claude 3 Opus outperformed Haiku in twoseparate testsResults published by Anthropic for MMLU and GPQA, percentage of correct answers, 2024.Made for Raúl Mata Meneses, by HanademiSources: Anthropic. (2024). The Claude 3 model family: Opus, Sonnet, Haiku.UnitpercentageMMLU75.2%79%86.8%HaikuSonnetOpusGPQA33.3%40.4%50.4%HaikuSonnetOpus
Haiku, Sonnet, and Opus belong to the same family but achieved different results. Opus led both MMLU and GPQA in the figures published by Anthropic. Choosing the brand without specifying the model leaves out an important difference.
Claude 3 Opus superó a Haiku en dospruebas distintasResultados publicados por Anthropic para MMLU y GPQA, porcentaje de respuestas correctas, 2024.Made for Raúl Mata Meneses, by HanademiFuentes: Anthropic. (2024). The Claude 3 model family: Opus, Sonnet, Haiku.UnidadporcentajeMMLU75,2 %79 %86,8 %HaikuSonnetOpusGPQA33,3 %40,4 %50,4 %HaikuSonnetOpus
Haiku, Sonnet y Opus pertenecen a la misma familia, pero obtuvieron resultados distintos. Opus lideró tanto MMLU como GPQA en las cifras publicadas por Anthropic. Elegir la marca sin especificar el modelo deja fuera una diferencia importante.
Equal Claude context windows hide 36 to 42point benchmark gapsMade for Raúl Mata Meneses, by HanademiSources: Anthropic. (2024). The Claude 3 model family: Opus, Sonnet, Haiku.; AI Model Benchmark Comparison.;Anthropic. (2024). Introducing the next generation of Claude.Unitpercentage score and percentage-point gapAll three models: 200,000-token context windowClaude 3 HaikuGPQA 33.3%MMLU 75.2%41.9-pt gapClaude 3 SonnetGPQA 40.4%MMLU 79.0%38.6-pt gapClaude 3 OpusGPQA 50.4%MMLU 86.8%36.4-pt gap0255075100
Opus scores higher on both tests and narrows the gap to 36.4 points, while Haiku's gap reaches 41.9 points despite the identical context capacity.
Ventanas de contexto iguales en Claude ocultanbrechas de 36 a 42 puntos entre pruebasMade for Raúl Mata Meneses, by HanademiFuentes: Anthropic. (2024). The Claude 3 model family: Opus, Sonnet, Haiku.; AI Model Benchmark Comparison.;Anthropic. (2024). Introducing the next generation of Claude.Unidadpuntuación porcentual y brecha en puntos porcentualesLos tres modelos: ventana de contexto de 200.000 tokensClaude 3 HaikuGPQA 33,3%MMLU 75,2%Brecha 41,9 ptClaude 3 SonnetGPQA 40,4%MMLU 79,0%Brecha 38,6 ptClaude 3 OpusGPQA 50,4%MMLU 86,8%Brecha 36,4 pt0255075100
Opus obtiene mejores resultados en ambas pruebas y reduce la brecha a 36,4 puntos, mientras la brecha de Haiku llega a 41,9 puntos pese a tener la misma capacidad de contexto.
A wide window admits more material, butdoes not guarantee reading it perfectly.The figure comes from Anthropic's announcement and could not be verified against a livesource during compilation.Sources: Anthropic. (2024). Introducing the next generation of Claude.
Anthropic announced the same nominal window of 200,000 tokens for all three Claude 3 models. That capacity allows inputting extensive documents or multiple sources. However, admitting text does not mean remembering, interpreting, or prioritizing each part with equal quality.
Una ventana amplia admite más material,pero no garantiza leerlo perfectamente.La cifra procede del anuncio de Anthropic y no pudo verificarse contra una fuente en vivodurante la compilación.Fuentes: Anthropic. (2024). Introducing the next generation of Claude.
Anthropic anunció la misma ventana nominal de 200.000 tokens para los tres modelos Claude 3. Esa capacidad permite introducir documentos extensos o varias fuentes. Sin embargo, admitir el texto no significa recordar, interpretar o priorizar cada parte con la misma calidad.
Gemini Ultra achieved 90% where Nanoreached 59.4%MMLU results published by Google DeepMind for Gemini Nano, Pro, and Ultra, 2023.Made for Raúl Mata Meneses, by HanademiSources: Google DeepMind. (2023). Gemini: A family of highly capable multimodal models.Unitpercentage90%Gemini Ultra79.1%Gemini Pro59.4%Gemini Nano
The Gemini family spans models designed for different needs. On MMLU, Google reported a wide gap between Nano and Ultra. That's why asking whether a task works with Gemini is not enough; you must identify the model, plan, and limits.
Gemini Ultra obtuvo 90% donde Nano alcanzó59,4%Resultados MMLU publicados por Google DeepMind para Gemini Nano, Pro y Ultra, 2023.Made for Raúl Mata Meneses, by HanademiFuentes: Google DeepMind. (2023). Gemini: A family of highly capable multimodal models.Unidadporcentaje90 %Gemini Ultra79,1 %Gemini Pro59,4 %Gemini Nano
La familia Gemini abarca modelos diseñados para necesidades distintas. En MMLU, Google reportó una distancia amplia entre Nano y Ultra. Por eso, preguntar si una tarea funciona con Gemini no basta; hay que identificar modelo, plan y límites.
Gemini expanded its nominal context to 2million tokensNominal windows published for Gemini 1.0 Pro and variants of Gemini 1.5, 2024. Logarithmic scale: eachline is x10.Made for Raúl Mata Meneses, by HanademiSources: Google. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.; Google for Developers. (2024).Gemini models.; Google Workspace. (2024). Gemini for Google Workspace pricing.UnittokensGemini 1.0 Pro32KGemini 1.5, 1 M1MGemini 1.5, 2 M2M
Google significantly expanded the amount of information that some Gemini variants could receive. The nominal jump was from 32,000 to 2 million tokens. That figure should be read alongside the version and plan, because the brand does not define a single experience.
Gemini amplió su contexto nominal hasta 2millones de tokensVentanas nominales publicadas para Gemini 1.0 Pro y variantes de Gemini 1.5, 2024.Made for Raúl Mata Meneses, by HanademiFuentes: Google. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.; Google for Developers. (2024).Gemini models.; Google Workspace. (2024). Gemini for Google Workspace pricing.UnidadtokensGemini 1.0 Pro32KGemini 1.5, 1 M1 MGemini 1.5, 2 M2 M
Google amplió fuertemente la cantidad de información que algunas variantes de Gemini podían recibir. El salto nominal fue desde 32.000 hasta 2 millones de tokens. Esa cifra debe leerse junto con la versión y el plan, porque la marca no define una experiencia única.
77% of early Copilot users did not want toabandon itUser perceptions of early Microsoft 365 Copilot adopters reported by Microsoft, 2023.Made for Raúl Mata Meneses, by HanademiSources: Microsoft. (2023). What can Copilot's earliest users teach us about generative AI at work?UnitpercentageGreater productivity70%Better quality68%Did not want to leave it77%
Early users expressed favorable signals about productivity, quality, and desire to keep Copilot. In Microsoft-sponsored tests, another average result indicated tasks completed 29% faster. These are useful signals to form a hypothesis, not a guarantee for each process.
El 77% de usuarios tempranos de Copilot noquería abandonarloPercepciones de usuarios tempranos de Microsoft 365 Copilot reportadas por Microsoft, 2023.Made for Raúl Mata Meneses, by HanademiFuentes: Microsoft. (2023). What can Copilot's earliest users teach us about generative AI at work?UnidadporcentajeMayor productividad70 %Mejor calidad68 %No quería dejarlo77 %
Los primeros usuarios expresaron señales favorables sobre productividad, calidad y deseo de conservar Copilot. En pruebas patrocinadas por Microsoft, otro resultado promedio indicó tareas completadas 29% más rápido. Son señales útiles para formular una hipótesis, no una garantía para cada proceso.
More than 60% of AI search citations wereincorrectThe Tow Center evaluated 8 AI search engines with 200 news fragments, producing 1,600 responses in2025. The chart uses a logarithmic scale.Made for Raúl Mata Meneses, by HanademiSources: Tow Center for Digital Journalism. (2025). AI search has a citation problem.; AI search engines fail to produce accuratecitations in over ....; AI Search Has a Citation Problem - Columbia Journalism Review.Unitcount / recuentoAI search engines / Buscadores con IA8News fragments / Fragmentos periodísticos200Responses / Respuestas1,600
The Tow Center examined 8 AI search engines, 200 news fragments, and 1,600 responses. Over 60% of citations were incorrect. These tools can accelerate discovery, but a visible link does not prove the source backs the answer.
Más de 60% de las citas de búsqueda con IAfueron incorrectasEl Tow Center evaluó 8 buscadores con IA con 200 fragmentos periodísticos, produciendo 1.600respuestas en 2025. El gráfico usa una escala logarítmica.Made for Raúl Mata Meneses, by HanademiFuentes: Tow Center for Digital Journalism. (2025). AI search has a citation problem.; AI search engines fail to produce accuratecitations in over ....; AI Search Has a Citation Problem - Columbia Journalism Review.Unidadcount / recuentoAI search engines / Buscadores con IA8News fragments / Fragmentos periodísticos200Responses / Respuestas1.600
El Tow Center examinó 8 buscadores con IA, 200 fragmentos periodísticos y 1.600 respuestas. Más de 60% de las citas fueron incorrectas. Estas herramientas pueden acelerar el descubrimiento, pero el enlace visible no demuestra que la fuente respalde la respuesta.
At Grok's error rate, only 12 of 200answers would be correctMade for Raúl Mata Meneses, by HanademiSources: Tow Center for Digital Journalism. (2025). AI search has a citation problem.; AI search engines fail toproduce accurate citations in over ....Unitanswers per 200 tested excerptsImplied result for each set of 200 answersPerplexity126 correct74 incorrectChatGPT Search66 correct134 incorrectGrok12 correct188 incorrect
The implied correct count falls from 126 for Perplexity to 66 for ChatGPT Search and 12 for Grok, making source verification essential.
Con la tasa de error de Grok, solo 12 de200 respuestas serían correctasMade for Raúl Mata Meneses, by HanademiFuentes: Tow Center for Digital Journalism. (2025). AI search has a citation problem.; AI search engines fail toproduce accurate citations in over ....Unidadrespuestas por cada 200 fragmentos evaluadosResultado implícito para cada conjunto de 200 respuestasPerplexity126 correctas74 incorrectasChatGPT Search66 correctas134 incorrectasGrok12 correctas188 incorrectas
La cantidad correcta implícita cae de 126 para Perplexity a 66 para ChatGPT Search y 12 para Grok, por lo que verificar las fuentes es esencial.
In one test, estimated error reached 94% inGrok.Approximate percentage of incorrect responses for three search engines in the Tow Center test, 2025;figures not verified live.Made for Raúl Mata Meneses, by HanademiSources: Tow Center for Digital Journalism. (2025). AI search has a citation problem.Unitpercentage94%Grok67%ChatGPT Search37%Perplexity
The test found weak results across all products tested, though with wide differences. Perplexity registered approximately 37%, ChatGPT Search 67%, and Grok 94% incorrect responses. These figures should be read as a specific test, not as a permanent ranking.
En una prueba, el error estimado llegó a94% en GrokPorcentaje aproximado de respuestas incorrectas para tres buscadores en la prueba del Tow Center, 2025;cifras no verificadas en vivo.Made for Raúl Mata Meneses, by HanademiFuentes: Tow Center for Digital Journalism. (2025). AI search has a citation problem.Unidadporcentaje94 %Grok67 %ChatGPT Search37 %Perplexity
La prueba encontró resultados débiles en todos los productos señalados, aunque con diferencias amplias. Perplexity registró cerca de 37%, ChatGPT Search 67% y Grok 94% de respuestas incorrectas. Estas cifras deben leerse como una prueba concreta, no como un ranking permanente.
Google disseminated the favorable evidence,so the company must validate it internally.Google reported that 75% of daily users perceived better quality. The 88% figure on speedrequires confirming the literal survey text.Sources: Google Workspace. (2024). The ROI of generative AI in the workplace.
The research commissioned by Google reported 105 minutes of average weekly savings and favorable quality signals. Those results help design a test, but the provider also has commercial interest in the outcome. The decision should be supported by your own tasks and data.
Google difundió la evidenciafavorable, así que la empresa debevalidarla internamente.Google reportó que 75% de usuarios diarios percibía mejor calidad. La cifra de 88%sobre rapidez requiere confirmar el texto literal del cuestionario.Fuentes: Google Workspace. (2024). The ROI of generative AI in the workplace.
La investigación encargada por Google reportó 105 minutos semanales de ahorro promedio y señales favorables de calidad. Esos resultados ayudan a diseñar una prueba, pero el proveedor también tiene interés comercial en el resultado. La decisión debe apoyarse en tareas y datos propios.
USD 20 per month converts to USD 24,000for 100 people.Annual cost before taxes of a USD 20 monthly subscription, by team size. Logarithmic scale: each line is×10.Made for Raúl Mata Meneses, by HanademiSources: OpenAI. (2023). Introducing ChatGPT Plus.; Microsoft. (2023). What can Copilot's earliest users teach us about generative AIat work?; National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: GenerativeArtificial Intelligence Profile.UnitUSD annually1 user$24010 users$2,400100 users$24,000
A personal subscription looks small until multiplied across your entire team. At USD 20 monthly, 100 people cost USD 24,000 per year before taxes. The purchase requires a frequent task, measurable savings, and review factored into the calculation.
USD 20 al mes se convierten en USD24.000 para 100 personasCosto anual antes de impuestos de una suscripción de USD 20 mensuales, por tamaño del equipo.Made for Raúl Mata Meneses, by HanademiFuentes: OpenAI. (2023). Introducing ChatGPT Plus.; Microsoft. (2023). What can Copilot's earliest users teach us about generative AI atwork?; National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative ArtificialIntelligence Profile.UnidadUSD anuales1 usuario$24010 usuarios$2.400100 usuarios$24.000
Una suscripción personal parece pequeña hasta que se multiplica por todo el equipo. A USD 20 mensuales, 100 personas cuestan USD 24.000 al año antes de impuestos. La compra necesita una tarea frecuente, ahorro medible y revisión incluida en el cálculo.
78% brought their own AI to work.Behaviors reported by people using AI at work, Microsoft and LinkedIn survey, 2024.Made for Raúl Mata Meneses, by HanademiSources: Microsoft & LinkedIn. (2024). AI at work is here. Now comes the hard part: 2024 Work Trend Index Annual Report.UnitpercentageBrings own AI78%Hides their use52%Fears being replaceable53%
Adoption is already happening even if the organization has not designed it. 78% of those using work AI reported bringing their own tools. The most practical response is to offer an authorized route that competes with the convenience of outside options.
El 78% llevaba su propia IA al trabajoConductas declaradas por personas que usaban IA en el trabajo, encuesta de Microsoft y LinkedIn, 2024.Made for Raúl Mata Meneses, by HanademiFuentes: Microsoft & LinkedIn. (2024). AI at work is here. Now comes the hard part: 2024 Work Trend Index Annual Report.UnidadporcentajeLleva IA propia78 %Oculta su uso52 %Teme ser reemplazable53 %
La adopción ya ocurre aunque la organización no la haya diseñado. El 78% de quienes usaban IA laboral declaró llevar herramientas propias. La respuesta más práctica es ofrecer una ruta autorizada que compita con la comodidad de las opciones externas.
Ease of use can turn a quick query intodata exposure.The figure comes from Cisco and has low confidence in this dataset. The breakdown bydata type requires further verification.Sources: Cisco. (2024). Cisco 2024 Data Privacy Benchmark Study.
Cisco reported that 48% of users had inputted non-public business information into generative tools. The risk appears before evaluating response quality. The first decision must be what information can enter and into which authorized tool.
La facilidad de uso puede convertir unaconsulta rápida en exposición de datos.La cifra procede de Cisco y tiene confianza baja en este conjunto. El desglose por tipo dedato requiere verificación adicional.Fuentes: Cisco. (2024). Cisco 2024 Data Privacy Benchmark Study.
Cisco reportó que 48% de los usuarios había introducido información empresarial no pública en herramientas generativas. El riesgo aparece antes de evaluar la calidad de la respuesta. La primera decisión debe ser qué información puede entrar y en qué herramienta autorizada.
Test the same tasks internally, and treatpreference rankings as one input ratherthan a final decision.Internal testing should cover four dimensions: quality, total time, data risk, and full cost.Preference rankings provide context, but their one methodological limitation is that theychange with models, votes, and methodology.Sources: National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework, AI RMF 1.0.; Measure - AIRC.;LMSYS Org. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference.; Chatbot Arena: An Open Platform for Evaluating LLMsby Human Preference.
The decision does not start with a brand or general ranking. It starts with representative tasks and criteria defined before testing. Use four dimensions for the internal test, and treat preference rankings as context rather than a substitute for process evaluation.
Prueba las mismas tareas internamente ytrata los rankings de preferencia como uninsumo, no como una decisión final.Las pruebas internas deben cubrir cuatro dimensiones: calidad, tiempo total, riesgo dedatos y costo completo. Los rankings de preferencia aportan contexto, pero tienen unalimitación metodológica: cambian con los modelos, los votos y la metodología.Fuentes: National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework, AI RMF 1.0.; Measure - AIRC.;LMSYS Org. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference.; Chatbot Arena: An Open Platform for Evaluating LLMsby Human Preference.
La decisión no empieza por una marca ni por un ranking general. Empieza con tareas representativas y criterios definidos antes de probar. Usa cuatro dimensiones en la prueba interna y trata los rankings de preferencia como contexto, no como sustituto de la evaluación del proceso.
In summaryMade for Raúl Mata Meneses, by HanademiSources: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School.; Tow Centerfor Digital Journalism. (2025). AI search has a citation problem.; Google. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.; Microsoft & LinkedIn. (2024). AI at work is here.Now comes the hard part: 2024 Work Trend Index Annual Report.; OpenAI. (2023). GPT-4 API general availability and deprecation of older models in the Completions API.; Reuters. (2024). OpenAI says ChatGPT's…En una tarea fuera de la frontera, los consultores con IA fueron aproximadamente 19 puntos porcentuales menospropensos a producir la respuesta correcta.El Tow Center evaluó 1.600 respuestas de ocho buscadores con IA y reportó que más de 60% de sus citas fueron incorrectas.Google amplió el contexto nominal desde 32.000 tokens en Gemini 1.0 Pro hasta uno y dos millones en variantes de Gemini 1.5.En una encuesta de 2024, 78% de quienes usaban IA laboral llevaba herramientas propias no proporcionadas formalmente por su organización.El precio de entrada de API cayó de 30 dólares en GPT-4 a 10 en GPT-4 Turbo y 5 en GPT-4o por millón de tokens.OpenAI informó que ChatGPT pasó de 200 millones de usuarios semanales en agosto a 300 millones en diciembre de 2024.
The value of the research is not only what each source knew, but what became visible when their evidence was combined.
En resumenMade for Raúl Mata Meneses, by HanademiFuentes: Dell'Acqua, F., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School.; Tow Centerfor Digital Journalism. (2025). AI search has a citation problem.; Google. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.; Microsoft & LinkedIn. (2024). AI at work is here.Now comes the hard part: 2024 Work Trend Index Annual Report.; OpenAI. (2023). GPT-4 API general availability and deprecation of older models in the Completions API.; Reuters. (2024). OpenAI says ChatGPT's…En una tarea fuera de la frontera, los consultores con IA fueron aproximadamente 19 puntos porcentuales menospropensos a producir la respuesta correcta.El Tow Center evaluó 1.600 respuestas de ocho buscadores con IA y reportó que más de 60% de sus citas fueron incorrectas.Google amplió el contexto nominal desde 32.000 tokens en Gemini 1.0 Pro hasta uno y dos millones en variantes de Gemini 1.5.En una encuesta de 2024, 78% de quienes usaban IA laboral llevaba herramientas propias no proporcionadas formalmente por su organización.El precio de entrada de API cayó de 30 dólares en GPT-4 a 10 en GPT-4 Turbo y 5 en GPT-4o por millón de tokens.OpenAI informó que ChatGPT pasó de 200 millones de usuarios semanales en agosto a 300 millones en diciembre de 2024.
El valor de la investigación no está solo en cada fuente, sino en lo que apareció al combinar sus evidencias.

The research behind this deck

AI was most valuable when the task matched its capabilities. A brand, a model, and an application are not the same thing.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

La IA fue más valiosa cuando la tarea coincidía con sus capacidades. Una marca, un modelo y una aplicación no son lo mismo.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English

Related researchInvestigación relacionada