Hanademi

The hidden bill for weak software

15 slides · 15 min · 1 minute ago language
ENES
theme
LightDark
view
DeckTableTalk
brand
HanademiPlatzi
The hidden bill for weaksoftwareMade for Jonathan Vivero, by Hanademi
La factura oculta del softwaredeficienteMade for Jonathan Vivero, by Hanademi
CISQ estimated a $2.41T software-qualitybill in 2022Estimated US cost of poor software quality, USD trillions, CISQ reports for 2018, 2020, and 2022.Made for Jonathan Vivero, by HanademiSources: Consortium for Information & Software Quality. (2022). The cost of poor software quality in the US: A 2022 report.UnitUSD trillions$2.0$2.5201820202022The 2022 estimate was $2.41trillion.$2.8
Software quality is not a housekeeping issue. CISQ estimated a $2.41 trillion US cost in 2022, after $2.84 trillion in 2018 and $2.08 trillion in 2020. The same report estimated accumulated technical debt at about $1.52 trillion. These are broad economic estimates, not a price tag for skipped testing alone.
CISQ estimó una factura de 2,41 billonespor mala calidad en 2022Coste estimado de la mala calidad del software en Estados Unidos, billones de dólares, informes CISQ de2018, 2020 y 2022.Made for Jonathan Vivero, by HanademiFuentes: Consortium for Information & Software Quality. (2022). The cost of poor software quality in the US: A 2022 report.Unidadbillones de dólares$2,0$2,5201820202022La estimación de 2022 fue de 2,41billones de dólares.$2,8
La calidad del software no es una tarea menor de mantenimiento. CISQ estimó un coste de 2,41 billones de dólares en Estados Unidos en 2022, después de 2,84 billones en 2018 y 2,08 billones en 2020. El mismo informe estimó la deuda técnica acumulada en unos 1,52 billones. Son estimaciones económicas amplias, no el precio aislado de omitir testing.
Six terms that matterMade for Jonathan Vivero, by HanademiTechnical debtFuture work created by choosing a faster or weaker solution.Unit testAn automated check of one small piece of software.Integration testA check that multiple components work correctly together.TDDTest-driven development writes a failing test before production code.Flaky testA test that passes or fails without a relevant code change.Code coverageThe share of code executed while tests run.
The debate often mixes different problems. Unit tests examine small pieces, integration tests examine interactions, and coverage only records what ran. TDD changes the development sequence. Technical debt describes the future burden left behind.
Seis términos que importanMade for Jonathan Vivero, by HanademiDeuda técnicaTrabajo futuro creado al elegir una solución más rápida o débil.Prueba unitariaUna comprobación automatizada de una pequeña parte del software.Prueba de integraciónUna comprobación de que varios componentes funcionan bien juntos.TDDEl desarrollo guiado por pruebas escribe una prueba fallida antes del código.Prueba flakyUna prueba que pasa o falla sin un cambio relevante de código.Cobertura de códigoLa parte del código ejecutada mientras corren las pruebas.
El debate suele mezclar problemas diferentes. Las pruebas unitarias examinan piezas pequeñas, las de integración examinan interacciones y la cobertura solo registra qué se ejecutó. TDD cambia la secuencia de desarrollo. La deuda técnica describe la carga futura que queda.
Foundations must be sharedA presenter speaks with a microphone on a blue-lit stage.Sources: Integration Testing: Checking How Components Work .... richard-seidl.com.
A common vocabulary is only useful when teams turn software fundamentals into shared practice.
Los fundamentos deben compartirseUn presentador habla con micrófono sobre un escenario iluminado en azul.Fuentes: Integration Testing: Checking How Components Work .... richard-seidl.com.
Un vocabulario común solo es útil cuando los equipos convierten los fundamentos del software en una práctica compartida.
NIST found $22.2B of testing costs might beavoidableEstimated 2002 US cost of inadequate software-testing infrastructure, USD billions.Made for Jonathan Vivero, by HanademiSources: Tassey, G. (2002). The economic impacts of inadequate infrastructure for software testing. National Institute of Standards andTechnology.Residual cost is the permitted remainder calculation: $59.5B minus $22.2B equals $37.3B.UnitUSD billions$0$20$40$60$59.5Total cost-$22.2Potentially avoidable$37.3Residual cost
NIST put a narrower number on inadequate testing infrastructure. Its 2002 estimate was $59.5 billion, with $22.2 billion potentially avoidable through better infrastructure. That leaves a large residual even inside this model. Testing matters, but it is not the whole quality system.
NIST estimó que 22.200 millones en costesde testing podrían evitarseCoste estimado en 2002 de una infraestructura de testing inadecuada en Estados Unidos, miles de millonesde dólares.Made for Jonathan Vivero, by HanademiFuentes: Tassey, G. (2002). The economic impacts of inadequate infrastructure for software testing. National Instituteof Standards and Technology.El coste residual es el cálculo permitido del remanente: 59.500 millones menos 22.200 millones equivalen a 37.300millones.Unidadmiles de millones de dólares$0$20$40$60$59,5Coste total-$22,2Potencialmente evitable$37,3Coste residual
NIST puso una cifra más estrecha a la infraestructura de testing inadecuada. Su estimación de 2002 fue de 59.500 millones de dólares, con 22.200 millones potencialmente evitables mediante una mejor infraestructura. Incluso dentro de este modelo queda un gran coste residual. El testing importa, pero no es todo el sistema de calidad.
The industrial TDD teams reported 40% to 90%fewer pre-release defects and 15% to 35% moreinitial time. Separately, a 32-developer MicrosoftNUnit team reported a 20.9% decrease in testdefect density.The table keeps the industrial TDD findings separate from the NUnit result because theycome from distinct source loci. The industrial results cover four teams and do notestablish a universal effect. The NUnit result supports automated unit testing withoutproving that tests must always come first.Sources: Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: Results andexperiences of four industrial teams. Empirical Software Engineering, 13, 289-302.; results and experiences of four industrial teams.; doi.org.
Four IBM and Microsoft teams reported 40% to 90% lower pre-release defect density and about 15% to 35% more initial development time after adopting TDD. Separately, a 32-developer Microsoft team using NUnit after coding for one year reported a 20.9% decrease in test defect density. These results support possible benefits from automated testing while not proving a universal TDD effect or a mandatory test-first sequence.
Los equipos industriales con TDD informaron entre40% y 90% menos defectos antes del lanzamiento yentre 15% y 35% más de tiempo inicial. Porseparado, un equipo de Microsoft de 32desarrolladores con NUnit informó una reduccióndel 20,9% en la densidad de defectos de prueba.La tabla mantiene separados los hallazgos industriales sobre TDD y el resultado conNUnit porque proceden de fuentes distintas. Los resultados industriales abarcan cuatroequipos y no establecen un efecto universal. El resultado con NUnit respalda las pruebasunitarias automatizadas sin demostrar que las pruebas deban escribirse siempreprimero.Fuentes: Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: Results andexperiences of four industrial teams. Empirical Software Engineering, 13, 289-302.; results and experiences of four industrial teams.; doi.org.
Cuatro equipos de IBM y Microsoft informaron entre 40% y 90% menos densidad de defectos antes del lanzamiento y aproximadamente entre 15% y 35% más de tiempo inicial tras adoptar TDD. Por separado, un equipo de Microsoft de 32 desarrolladores que usó NUnit después de programar durante un año informó una reducción del 20,9% en la densidad de defectos de prueba. Estos resultados respaldan posibles beneficios de las pruebas automatizadas, pero no demuestran un efecto universal de TDD ni una secuencia test-first obligatoria.
Quality involves tradeoffsA black weight and several smaller weights rest on opposite pans of a brass balance.Sources: Coupling Should Be Weighed, Not Counted. vladikk.com.
The TDD evidence makes the bargain explicit: more initial effort can accompany fewer defects.
La calidad implica compensacionesUna pesa negra y varias pesas más pequeñas descansan en platos opuestos de una balanza de latón.Fuentes: Coupling Should Be Weighed, Not Counted. vladikk.com.
La evidencia sobre TDD hace explícito el intercambio: un mayor esfuerzo inicial puede acompañar a menos defectos.
Google's 70-20-10 mix is guidance, not alawGoogle's published testing mix: 70% unit, 20% integration, and 10% end-to-end tests.Made for Jonathan Vivero, by HanademiSources: Google. (2015). Just say no to more end-to-end tests. Google Testing Blog.; Knight, J. C., & Leveson, N. G. (1986). Anexperimental evaluation of the assumption of independence in multiversion programming. IEEE Transactions on Software Engineering,SE-12(1), 96-109.UnitShare of tests0%20%40%60%80%100%Unit testsIntegration testsEnd-to-end testsPublished guidance70%20%10%
Google's pyramid gives teams a useful bias toward fast unit tests. It does not establish a universal optimum. A classic experiment tested about one million inputs across 27 independently developed versions and still found correlated failures. Components can repeat the same mistake, so isolation is not enough.
La mezcla 70-20-10 de Google es unaguía, no una leyMezcla de testing publicada por Google: 70% unitarias, 20% de integración y 10% end-to-end.Made for Jonathan Vivero, by HanademiFuentes: Google. (2015). Just say no to more end-to-end tests. Google Testing Blog.; Knight, J. C., & Leveson, N. G. (1986). Anexperimental evaluation of the assumption of independence in multiversion programming. IEEE Transactions on SoftwareEngineering, SE-12(1), 96-109.UnidadProporción de pruebas0 %20 %40 %60 %80 %100 %Pruebas unitariasPruebas de integraciónPruebas end-to-endGuía publicada70 %20 %10 %
La pirámide de Google ofrece a los equipos una preferencia útil por pruebas unitarias rápidas. No establece un óptimo universal. Un experimento clásico probó cerca de un millón de entradas en 27 versiones desarrolladas de forma independiente y encontró fallos correlacionados. Los componentes pueden repetir el mismo error, por lo que probarlos de forma aislada no basta.
31,000 test suites showed why coveragealone cannot certify effective testing.Coverage is a useful gap detector, but a weak standalone quality target. A guard can visitevery room without noticing every open window.Sources: Inozemtseva, L., & Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. International Conference on SoftwareEngineering.
Coverage answers a narrow question: did the test execute this code? Across more than 31,000 suites from five projects, coverage was only weakly associated with effectiveness after controlling for suite size. It still exposes completely untested code. What it cannot provide is a universal percentage that guarantees meaningful fault detection.
31.000 suites mostraron por qué lacobertura sola no certifica un testing eficaz.La cobertura es útil para detectar huecos, pero es un objetivo de calidad débil por sísolo. Un vigilante puede visitar todas las habitaciones sin detectar cada ventana abierta.Fuentes: Inozemtseva, L., & Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. International Conference on SoftwareEngineering.
La cobertura responde a una pregunta estrecha: ¿la prueba ejecutó este código? En más de 31.000 suites de cinco proyectos, la cobertura solo se asoció débilmente con la eficacia tras controlar el tamaño de la suite. Sigue sirviendo para descubrir código sin probar. Lo que no ofrece es un porcentaje universal que garantice la detección de fallos significativos.
Flaky-test noise is documented throughthree separate evidence points.The evidence comes from different sources and measures different things: a projectsample, Google's execution rate, and a study of fix commits.Sources: Hilton, M., Tunnell, T., Huang, K., Marinov, D., & Dig, D. (2016). Usage, costs, and benefits of continuous integration in open-source projects. Automated Software EngineeringConference.; Usage, Costs, and Benefits of Continuous Integration in Open ....; Micco, J. (2016). Flaky tests at Google and how we mitigate them. Google Testing Blog.; Flaky Tests at Google.;Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). An empirical analysis of flaky tests. Foundations of Software Engineering.; An Empirical Analysis of Flaky Tests.
Continuous integration shortens the path from change to feedback, but flaky tests can make that feedback noisy. One study covered 34,544 open-source projects. Google reported that 1.5% of its test executions produced flaky outcomes. A separate empirical study examined 201 commits fixing flaky tests and identified recurring causes including asynchronous waits, concurrency, and test-order dependencies.
El ruido de las pruebas flaky se documentamediante tres evidencias separadas.La evidencia procede de fuentes distintas y mide cosas diferentes: una muestra deproyectos, la tasa de ejecuciones de Google y un estudio de commits de corrección.Fuentes: Hilton, M., Tunnell, T., Huang, K., Marinov, D., & Dig, D. (2016). Usage, costs, and benefits of continuous integration in open-source projects. Automated Software EngineeringConference.; Usage, Costs, and Benefits of Continuous Integration in Open ....; Micco, J. (2016). Flaky tests at Google and how we mitigate them. Google Testing Blog.; Flaky Tests at Google.;Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). An empirical analysis of flaky tests. Foundations of Software Engineering.; An Empirical Analysis of Flaky Tests.
La integración continua acorta el camino entre un cambio y el feedback, pero las pruebas flaky pueden volver ruidosa esa respuesta. Un estudio abarcó 34.544 proyectos abiertos. Google informó que el 1,5% de sus ejecuciones produjo resultados flaky. Otro estudio empírico examinó 201 commits que corregían pruebas flaky e identificó causas recurrentes, como esperas asíncronas, concurrencia y dependencias del orden de las pruebas.
Technical debt can consume 10% to 20% ofproduct budgetsShare of technology budgets diverted from new products, as reported by McKinsey. The supplied midpoint is15%.Made for Jonathan Vivero, by HanademiSources: McKinsey & Company. (2020). Tech debt: Reclaiming tech equity.; Demystifying digital dark matter: A new standard to tame technical debt.20%Upper bound / Límite superior15%Midpoint / Punto medio10%Lower bound / Límite inferior
Technical debt is not only an engineering complaint. McKinsey reported a 10% lower bound, a 15% supplied midpoint, and a 20% upper bound for technology budgets diverted from new products. The roadmap must absorb that displaced capacity.
La deuda técnica puede consumir del 10% al20% del presupuesto de productoProporción de presupuestos tecnológicos desviados de nuevos productos, según McKinsey. El punto medioaportado es 15%.Made for Jonathan Vivero, by HanademiFuentes: McKinsey & Company. (2020). Tech debt: Reclaiming tech equity.; Demystifying digital dark matter: A new standard to tame technical debt.20 %Upper bound / Límite superior15 %Midpoint / Punto medio10 %Lower bound / Límite inferior
La deuda técnica no es solo una queja de ingeniería. McKinsey informó de un límite inferior del 10%, un punto medio aportado del 15% y un límite superior del 20% para los presupuestos tecnológicos desviados de nuevos productos. La hoja de ruta debe absorber esa capacidad desplazada.
Synopsys found high-risk flaws in 74% ofaudited codebasesShare of codebases in Synopsys's 2024 audit sample containing open source, known vulnerabilities, andhigh-risk vulnerabilities.Made for Jonathan Vivero, by HanademiSources: Synopsys. (2024). Open Source Security and Risk Analysis report.UnitAudited codebasesContains open source · 96%Contains vulnerabilities · 84%Contains high-risk flaws · 74%
Modern software inherits code, maintenance obligations, and vulnerabilities from its dependencies. Synopsys reported open source in 96% of audited codebases, vulnerabilities in 84%, and high-risk vulnerabilities in 74%. Testing can expose some failures, but it cannot exhaust every dependency state and attacker behavior. Secure engineering must also govern what enters the system and how teams respond.
Synopsys halló fallos de alto riesgo en el74% de las bases auditadasProporción de bases en la muestra auditada por Synopsys en 2024 con código abierto, vulnerabilidadesconocidas y vulnerabilidades de alto riesgo.Made for Jonathan Vivero, by HanademiFuentes: Synopsys. (2024). Open Source Security and Risk Analysis report.UnidadBases de código auditadasContiene código abierto · 96 %Contiene vulnerabilidades · 84 %Contiene fallos de alto riesgo · 74 %
El software moderno hereda código, obligaciones de mantenimiento y vulnerabilidades de sus dependencias. Synopsys informó de código abierto en el 96% de las bases auditadas, vulnerabilidades en el 84% y vulnerabilidades de alto riesgo en el 74%. El testing puede revelar algunos fallos, pero no puede agotar cada estado de dependencias y comportamiento atacante. La ingeniería segura también debe gobernar qué entra en el sistema y cómo responde el equipo.
Teams can improve a target metric withoutimproving software.The operating rule is simple: use several imperfect checks together and keep customeroutcomes above activity metrics.Sources: Strathern, M. (1997). Improving ratings: Audit in the British university system. European Review.
Coverage, warning counts, and defect totals are signals, not definitions of quality. Once one signal becomes the target, teams can optimize the reported number while customer outcomes remain unchanged. The safer system combines small reviews, automated analysis, layered tests, dependency control, and production outcomes. No single score gets to declare victory.
Los equipos pueden mejorar una métricaobjetivo sin mejorar el software.La regla operativa es sencilla: usar varias comprobaciones imperfectas juntas y situarlos resultados del cliente por encima de las métricas de actividad.Fuentes: Strathern, M. (1997). Improving ratings: Audit in the British university system. European Review.
La cobertura, los avisos y los defectos son señales, no definiciones de calidad. Cuando una señal se convierte en la meta, los equipos pueden optimizar la cifra informada mientras los resultados del cliente no cambian. El sistema más seguro combina revisiones pequeñas, análisis automatizado, pruebas por capas, control de dependencias y resultados en producción. Ninguna puntuación puede declarar la victoria por sí sola.
Metrics cannot replace judgmentSource code is open on a laptop positioned beside a potted plant.Sources: Your Static Analysis Tool Is Lying to You About Complexity - Codequiry Blog | Codequiry. codequiry.com.
After governance checks the score, developers still make design and testing decisions in the code.
Las métricas no reemplazan el criterioEl código fuente está abierto en un portátil situado junto a una planta en maceta.Fuentes: Your Static Analysis Tool Is Lying to You About Complexity - Codequiry Blog | Codequiry. codequiry.com.
Después de revisar las métricas, los desarrolladores aún toman decisiones de diseño y testing en el código.
The economic burden is visible, butquality requires several defensesworking together.CISQ reported $2.84T in 2018, $2.08T in 2020, and $2.41T in 2022 for poor softwarequality. Other evidence shows possible TDD benefits, the limits of coverage, flaky-testnoise at Google, and high-risk vulnerabilities in 74% of audited codebases.Sources: Consortium for Information & Software Quality. (2022). The cost of poor software quality in the US: A 2022 report.; Cost of Poor Software Quality in the U.S.: A 2022 Report.; Nagappan, N., Maximilien, E. M.,Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: Results and experiences of four industrial teams. Empirical Software Engineering, 13, 289-302.; results and experiencesof four industrial teams.; Inozemtseva, L., & Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. International Conference on Software Engineering.; Coverage is not strongly correlated…
The economic exposure is enormous, but the evidence does not support one universal remedy. TDD may reduce pre-release defects, while coverage alone cannot certify effective testing. Flaky executions can pollute feedback, and audited codebases often contain serious dependency risks. Quality therefore depends on several defenses and on measuring outcomes, not one ritual or score.
La carga económica es visible,pero la calidad requiere variasdefensas coordinadas.CISQ informó de 2,84 billones de USD en 2018, 2,08 billones en 2020 y 2,41 billones en2022 por la mala calidad del software. Otros datos muestran posibles beneficios delTDD, los límites de la cobertura, el ruido de las pruebas flaky en Google yvulnerabilidades de alto riesgo en el 74% de las bases auditadas.Fuentes: Consortium for Information & Software Quality. (2022). The cost of poor software quality in the US: A 2022 report.; Cost of Poor Software Quality in the U.S.: A 2022 Report.; Nagappan, N., Maximilien, E. M.,Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: Results and experiences of four industrial teams. Empirical Software Engineering, 13, 289-302.; results and experiencesof four industrial teams.; Inozemtseva, L., & Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. International Conference on Software Engineering.; Coverage is not strongly correlated…
La exposición económica es enorme, pero los datos no respaldan una solución universal. El TDD puede reducir los defectos antes del lanzamiento, mientras que la cobertura por sí sola no certifica pruebas eficaces. Las ejecuciones flaky pueden contaminar el feedback y las bases auditadas suelen contener riesgos graves en sus dependencias. Por tanto, la calidad depende de varias defensas y de medir resultados, no de un único ritual o puntuación.
In summaryMade for Jonathan Vivero, by HanademiSources: Consortium for Information & Software Quality. (2022). The cost of poor software quality in the US: A 2022 report.; Synopsys. (2024). Open Source Security and Risk Analysis report.; Inozemtseva, L., &Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. International Conference on Software Engineering.; doi.org.; Knight, J. C., & Leveson, N. G. (1986). An experimental evaluation of theassumption of independence in multiversion programming. IEEE Transactions on Software Engineering, SE-12(1), 96-109.; McKinsey & Company. (2020). Tech debt: Reclaiming tech equity.CISQ estimated US poor-software-quality costs at $2.84 trillion in 2018, $2.08 trillion in 2020, and $2.41 trillion in 2022.Synopsys reported open source in 96% of audited codebases, vulnerabilities in 84%, and high-risk vulnerabilities in 74% in its 2024 sample.Across more than 31,000 test suites from five projects, coverage was only weakly associated with effectiveness aftercontrolling for test-suite size.A 32-developer Microsoft team using NUnit after coding for one year achieved a 20.9% decrease in test defect density.Testing about one million inputs across 27 independently developed program versions revealed correlated failures,contradicting an assumption of independent implementation errors.McKinsey reported that technical debt often diverted 10% to 20% of technology budgets intended for new products.
The value of the research is not only what each source knew, but what became visible when their evidence was combined.
En resumenMade for Jonathan Vivero, by HanademiFuentes: Consortium for Information & Software Quality. (2022). The cost of poor software quality in the US: A 2022 report.; Synopsys. (2024). Open Source Security and Risk Analysis report.; Inozemtseva, L., &Holmes, R. (2014). Coverage is not strongly correlated with test suite effectiveness. International Conference on Software Engineering.; doi.org.; Knight, J. C., & Leveson, N. G. (1986). An experimental evaluation of theassumption of independence in multiversion programming. IEEE Transactions on Software Engineering, SE-12(1), 96-109.; McKinsey & Company. (2020). Tech debt: Reclaiming tech equity.CISQ estimó el coste estadounidense del software de mala calidad en 2,84 billones de dólares en 2018, 2,08 billones en2020 y 2,41 billones en 2022.Synopsys informó de código abierto en el 96% de las bases auditadas, vulnerabilidades en el 84% y vulnerabilidades de altoriesgo en el 74% en 2024.En más de 31.000 suites de cinco proyectos, la cobertura mostró una asociación débil con la eficacia tras controlar el tamaño de la suite.Un equipo de Microsoft de 32 desarrolladores que usó NUnit después de programar logró una reducción de 20,9% en ladensidad de defectos de prueba tras un año.Probar cerca de un millón de entradas en 27 versiones independientes reveló fallos correlacionados, contradiciendo lasupuesta independencia de sus errores.McKinsey informó que la deuda técnica desviaba frecuentemente entre el 10% y el 20% de los presupuestos tecnológicosdestinados a nuevos productos.
El valor de la investigación no está solo en cada fuente, sino en lo que apareció al combinar sus evidencias.

The research behind this deck

The bill fell after 2018, then climbed back to $2.41 trillion. These terms separate test quantity, test scope, and long-term engineering cost.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

La factura cayó después de 2018 y volvió a subir hasta 2,41 billones de dólares. Estos términos separan cantidad de pruebas, alcance y coste de ingeniería a largo plazo.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English