Hanademi

The cost of ignoring dataEl costo de ignorar los datos · V8 Master Engine

30 slides · 19 min · 2 hours ago30 láminas · 19 min · hace 2 h language
ENES
theme
LightDark
view
TalkTable
brand
HanademiPlatzi
The cost of ignoring dataMade for Johannes Talero, by Hanademi
  1. The problem was not lack of data, but deciding after excluding cases that contradicted the convenient narrative.
  2. The O-ring operated near 31°F, 22°F outside the observed range before launch.
  3. When the cost of error is extreme, uncertainty must constrain deployment, not justify it.
El costo de ignorar los datosMade for Johannes Talero, by Hanademi
  1. El problema no fue carecer de datos, sino decidir después de excluir los casos que contradecían la historia conveniente.
  2. La junta operó cerca de 31 °F, 22 °F fuera del rango observado antes del lanzamiento.
  3. Cuando el costo del error es extremo, la incertidumbre debe frenar el despliegue, no justificarlo.
All 4 pre-Challenger flights below 65°F hadO-ring damageMade for Johannes Talero, by HanademiSources: American Statistical Association; American Association for the Advancement of Science; theregister.com; PMLR;media.mit.edu; American Medical Association; Zillow Group; apnews.comUnit% of each observed groupPre-Challenger flights below 65°F with O-ring damage100%Google Flu Trends weeks above CDC reports, 108-week review92.6%Sepsis patients without a timely alert, one hospital implementation67%Gender Shades maximum error for dark-skinned women34.7%Announced Zillow workforce reduction25%
Across five cases, the cost appeared in damaged flights, persistent overestimation, missed patients, unequal errors, and layoffs.
Los 4 vuelos previos al Challenger bajo 65°F tuvieron daños en juntas tóricasMade for Johannes Talero, by HanademiFuentes: American Statistical Association; American Association for the Advancement of Science; theregister.com;PMLR; media.mit.edu; American Medical Association; Zillow Group; apnews.comUnidad% de cada grupo observadoVuelos previos al Challenger bajo 65 °F con daños en juntas tóricas100%Semanas de Google Flu Trends por encima de los CDC, revisión de 108 semanas92,6%Pacientes con sepsis sin alerta oportuna, una implementación hospitalaria67%Error máximo de Gender Shades para mujeres de piel oscura34,7%Reducción de plantilla anunciada por Zillow25%
En cinco casos, el costo apareció como vuelos dañados, sobrestimación persistente, pacientes omitidos, errores desiguales y despidos.
The words that sustain the caseMade for Johannes Talero, by HanademiO-ringFlexible seal that closes a joint and prevents gas leakage.ExtrapolationPrediction beyond the range of data the model observed.SensitivityProportion of actual cases a system detects.Positive predictive valueProportion of alerts that correspond to actual cases.AI RMFNIST framework for managing artificial intelligence risks.
This case joins engineering, statistics, and organizational decisions. These definitions let you read the rest of the deck without confusing a prediction with an observation. The most important distinction is simple: extrapolation means leaving known ground.
Las palabras que sostienen el casoMade for Johannes Talero, by HanademiO-ringAnillo flexible que sella una unión e impide la fuga de gases.ExtrapolaciónPredicción fuera del rango de datos que el modelo observó.SensibilidadProporción de casos reales que un sistema detecta.Valor predictivo positivoProporción de alertas que corresponden a casos reales.AI RMFMarco de NIST para gestionar riesgos de inteligencia artificial.
Este caso une ingeniería, estadística y decisiones organizacionales. Estas definiciones permiten leer el resto del deck sin confundir una predicción con una observación. La distinción más importante es simple: extrapolar significa salir del terreno conocido.
At 53°F, the flight had 3 incidents ofrubber-seal damageDamage index and incident count at the 4 temperatures below 65°F documented in pre-Challenger records.Made for Johannes Talero, by HanademiSources: Presidential Commission on the Space Shuttle Challenger Accident. (1986). Report to the President, Volume I.U.S. Government Printing Office.; randomservices.org.; math.montana.edu.Values from 2 records were combined by temperature.Unitdamage index and incidents0123Incidentes25810Índice de daño →53°F57°F58°F63°FAt 53°F, the damage index was 11and there were 3 incidents.
You already know the ending: Challenger disintegrated 73 seconds after launch. The question is what information existed before. On cold flights, the 53°F point concentrated the highest damage index and 3 incidents, a warning hard to dismiss as noise.
A 53 °F, el vuelo tuvo 3 incidentes de dañoen sellos de gomaÍndice de daño y número de incidentes en las 4 temperaturas bajo 65 °F compartidas por los registrosprevios al Challenger.Made for Johannes Talero, by HanademiFuentes: Presidential Commission on the Space Shuttle Challenger Accident. (1986). Report to the President, Volume I.U.S. Government Printing Office.; randomservices.org.; math.montana.edu.Se unieron por temperatura los valores publicados en 2 registros.Unidadíndice de daño e incidentes0123Incidentes25810Índice de daño →53°F57°F58°F63°FA 53 °F, el índice fue 11 y hubo 3incidentes.
El lector ya conoce el final: el Challenger se desintegró 73 segundos después del despegue. La pregunta es qué información existía antes. En los vuelos fríos, el punto de 53 °F concentraba el mayor índice de daño y 3 incidentes, una advertencia difícil de llamar ruido.
The stakes leave the spreadsheetA space shuttle rises above a launch pad as exhaust and smoke spread below.Sources: Lessons from Challenger - Aerospace America - AIAA. aerospaceamerica.aiaa.org.
The January 28, 1986 launch turned incomplete analysis into an irreversible operational decision.
El riesgo sale de la hoja de cálculoUn transbordador espacial se eleva sobre una plataforma mientras el escape y el humo se extiendendebajo.Fuentes: Lessons from Challenger - Aerospace America - AIAA. aerospaceamerica.aiaa.org.
El lanzamiento del 28 de enero de 1986 convirtió un análisis incompleto en una decisión operativa irreversible.
Keeping only damaged flights made 17% and14% damage rates become 100%Made for Johannes Talero, by HanademiSources: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Airtransportation safety investigation report A20W0016 - Transportation Safety Board of Canada.; Dalal, S. R., Fowlkes, E. B., & Hoadley, B.(1989). Risk analysis of the space shuttle: Pre-Challenger prediction of failure. Journal of the American Statistical Association.UnitPercentage of flights with O-ring damageBelow 65°F100% in both samples65°F to 74°F16.7% with full data100% incident-only75°F or warmer14.3% with full data100% incident-only
Removing the 16 flights without damage preserved a 100% rate in every temperature band and erased the association between cold and damage.
Conservar solo vuelos dañados convirtiótasas de daño de 17% y 14% en 100%Made for Johannes Talero, by HanademiFuentes: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Airtransportation safety investigation report A20W0016 - Transportation Safety Board of Canada.; Dalal, S. R., Fowlkes, E. B., & Hoadley,B. (1989). Risk analysis of the space shuttle: Pre-Challenger prediction of failure. Journal of the American Statistical Association.UnidadPorcentaje de vuelos con daño en O-ringsMenos de 65°F100% en ambas muestrasDe 65°F a 74°F16.7% con datos completos100% solo incidentes75°F o más14.3% con datos completos100% solo incidentes
Eliminar los 16 vuelos sin daño conservó una tasa de 100% en cada rango de temperatura y borró la asociación entre frío y daño.
The O-ring operated 22°F outside theobserved rangeTemperatures of the prior minimum, launch environment, and estimated critical O-ring, in Fahrenheit.Made for Johannes Talero, by HanademiSources: National Aeronautics and Space Administration. (1986). Report of the Presidential Commission on the Space Shuttle Challenger Accident.NASA.; Presidential Commission on the Space Shuttle Challenger Accident. (1986). Report to the President, Volume I, Chapter 5. U.S. GovernmentPrinting Office.; journals.zeuspress.org.Unit°F0204053Prior minimum36Launch environment31Critical O-ringLimit recommended by engineers
The coldest prior flight had occurred at 53°F. The launch environment was near 36°F and the critical O-ring was estimated at 31°F. Engineers initially recommended not launching below 53°F, but the decision ultimately required proving danger in territory with no prior experience.
La junta operó 22 °F fuera del rangoobservadoTemperaturas del mínimo previo, del ambiente de lanzamiento y de la junta crítica estimada, en gradosFahrenheit.Made for Johannes Talero, by HanademiFuentes: National Aeronautics and Space Administration. (1986). Report of the Presidential Commission on the Space Shuttle Challenger Accident.NASA.; Presidential Commission on the Space Shuttle Challenger Accident. (1986). Report to the President, Volume I, Chapter 5. U.S. GovernmentPrinting Office.; journals.zeuspress.org.Unidad°F0204053Mínimo previo36Ambiente del lanzamiento31Junta críticaLímite recomendado por ingenieros
El vuelo previo más frío había ocurrido a 53 °F. El ambiente del lanzamiento estaba cerca de 36 °F y la junta crítica se estimó en 31 °F. Los ingenieros recomendaron inicialmente no lanzar por debajo de 53 °F, pero la decisión terminó exigiendo demostrar peligro en un terreno sin experiencia previa.
A later model put the damage chance above50% at 31°FMade for Johannes Talero, by HanademiSources: National Aeronautics and Space Administration. (1986). Report of the Presidential Commission on the Space ShuttleChallenger Accident. NASA.; v1ch3.; UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset.University of California, Irvine.UnitDegrees FahrenheitEstimated Challenger joint temperature31°FColdest prior flight53°FModeled temperature for 50% damage probability63.45°F
The launch condition was not merely colder than experience: it lay far below the temperature where the later model placed even odds of damage.
Un modelo posterior estimó más de 50% deprobabilidad de daño a 31 °FMade for Johannes Talero, by HanademiFuentes: National Aeronautics and Space Administration. (1986). Report of the Presidential Commission on the Space ShuttleChallenger Accident. NASA.; v1ch3.; UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset.University of California, Irvine.UnidadGrados FahrenheitTemperatura estimada de la junta del Challenger31°FVuelo previo más frío53°FTemperatura modelada para 50% de probabilidad de daño63.45°F
La condición de lanzamiento no solo fue más fría que la experiencia: quedó muy por debajo de la temperatura donde el modelo posterior situó el daño como igualmente probable.
Waiting until 60°F could cut estimatedshuttle-loss risk from 13% to 2%Estimated minimum risk of catastrophic field joint failure across 23 prior launches.Made for Johannes Talero, by HanademiSources: sfu.ca.Unit% probability31°F13%60°F2%6.5x
The dilemma was not launch or achieve zero risk. The available comparison was between at least 13% at 31 °F and at least 2% at 60 °F. Waiting materially changed the exposure without requiring perfect prediction.
Sources
Esperar a 60 °F reducía el riesgo estimadode perder el transbordador del 13% al 2%Riesgo mínimo estimado de falla catastrófica de la junta de campo con 23 lanzamientos previos.Made for Johannes Talero, by HanademiFuentes: sfu.ca.Unidad% de probabilidad31°F13 %60°F2 %6,5x
El dilema no era lanzar o alcanzar riesgo cero. La comparación disponible era entre al menos 13% a 31 °F y al menos 2% a 60 °F. Esperar cambiaba materialmente la exposición sin exigir una predicción perfecta.
Fuentes
Managers estimated disaster 13,000 timesless likely than seal damage at 31°FConvert 13%, 1 in 100 and 1 in 100,000 to probabilities, then divide each by management's estimate.Made for Johannes Talero, by HanademiSources: sfu.ca.; nasa.gov.UnitMultiple of management's estimated loss probabilityData estimate at 31°F13,000Working engineers1,000Management1
The data-based estimate exceeded the working-engineer estimate by about 13 times and the management estimate by about 13,000 times.
Directivos estimaron el desastre 13.000 vecesmenos probable que el daño del sello a 31 °FConvertir 13%, 1 de cada 100 y 1 de cada 100,000 en probabilidades y dividir cada una entre la estimaciónde la dirección.Made for Johannes Talero, by HanademiFuentes: sfu.ca.; nasa.gov.UnidadMúltiplo de la probabilidad de pérdida estimada por la direcciónEstimación con datos a 31°F13.000Ingenieros1.000Dirección1
La estimación basada en datos superó unas 13 veces la de los ingenieros y unas 13,000 veces la de la dirección.
The 23-flight model predicts 0.82probability of damage at 31 °F.The estimate depends on model specification and was not available during theteleconference. Its value lies in showing that the complete record contained a quantifiablesignal, not in claiming certainty beyond the observed range.Sources: utstat.toronto.edu.
The complete record allows estimation of a probability that the incident sample alone cannot identify. The shuttle2 model predicts 0.82 probability of damage at 31 °F. The figure was calculated after the accident and extrapolates 22 °F below the observed minimum.
El modelo de 23 vuelos predice 0.82 deprobabilidad de daño a 31 °F.La estimación depende de la especificación del modelo y no estuvo disponible durante lateleconferencia. Su valor está en mostrar que el registro completo contenía una señalcuantificable, no en prometer certeza fuera del rango observado.Fuentes: utstat.toronto.edu.
El registro completo permite estimar una probabilidad que la muestra de incidentes no puede identificar. El modelo shuttle2 predice 0.82 de probabilidad de daño a 31 °F. La cifra se calculó después del accidente y extrapola 22 °F bajo el mínimo observado.
All 4 flights below 65°F had damagedrubber sealsPrior flights with O-ring thermal damage, grouped by launch temperature.Made for Johannes Talero, by HanademiSources: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California,Irvine.; Dalal, S. R., Fowlkes, E. B., & Hoadley, B. (1989). Risk analysis of the space shuttle: Pre-Challenger prediction of failure.Journal of the American Statistical Association.Unitflights with damage4Below 65 °F265 to 74 °F175 °F or above
All 4 flights below 65 °F showed damage. Above that threshold there were also 3 damaged flights, so temperature was not a perfect rule. The error was demanding perfection from a risk signal before allowing it to influence the decision.
Los 4 vuelos bajo 65 °F tuvieron sellos degoma dañadosVuelos previos con daño térmico de O-rings, agrupados por temperatura de lanzamiento.Made for Johannes Talero, by HanademiFuentes: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.;Dalal, S. R., Fowlkes, E. B., & Hoadley, B. (1989). Risk analysis of the space shuttle: Pre-Challenger prediction of failure. Journal ofthe American Statistical Association.Unidadvuelos con daño4Menos de 65 °F265 a 74 °F175 °F o más
Los 4 vuelos bajo 65 °F presentaron daño. Sobre ese umbral también hubo 3 vuelos dañados, así que la temperatura no era una regla perfecta. El error fue exigir perfección a una señal de riesgo antes de permitir que influyera en la decisión.
Flights above 65°F had 4 incidents ofdamage to rubber sealsPrimary erosion or blowby incidents on prior flights above 65 °F.Made for Johannes Talero, by HanademiSources: math.montana.edu.Unitincidents70°F170°F175°F2
Warmer flights were not free of damage. The record totals 4 incidents above 65 °F. That prevents a simple narrative, but does not erase the concentration of severity at the coldest temperatures.
Los vuelos sobre 65 °F tuvieron 4incidentes de daño en sellos de gomaIncidentes primarios de erosión o soplado en vuelos previos por encima de 65 °F.Made for Johannes Talero, by HanademiFuentes: math.montana.edu.Unidadincidentes70°F170°F175°F2
Los vuelos más cálidos no estuvieron libres de daño. El registro suma 4 incidentes por encima de 65 °F. Eso impide una historia simplista, pero no borra la concentración de gravedad en las temperaturas más bajas.
The presentation omitted 16 of 23 flightsTotal prior flights, flights with damage, and flights without damage left out of the incident-focused argument.Made for Johannes Talero, by HanademiSources: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Mangel, M.,& Samaniego, F. J. (1984). Abraham Wald's work on aircraft survivability. Journal of the American Statistical Association.; cna.org.Unitflights23Complete record16Without damage omitted7With damage
The record held 23 flights, but the simplified argument focused on 7 with damage. The 16 undamaged flights were neither noise nor irrelevant cases. They were the denominator needed to ask how risk changed with temperature.
La presentación dejó fuera 16 de 23 vuelosVuelos previos totales, vuelos con daño y vuelos sin daño omitidos del argumento centrado en incidentes.Made for Johannes Talero, by HanademiFuentes: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Mangel, M.,& Samaniego, F. J. (1984). Abraham Wald's work on aircraft survivability. Journal of the American Statistical Association.; cna.org.Unidadvuelos23Registro completo16Sin daño omitidos7Con daño
El registro tenía 23 vuelos, pero el argumento simplificado se concentró en 7 con daño. Los 16 vuelos sin daño no eran ruido ni casos irrelevantes. Eran el denominador necesario para preguntar cómo cambiaba el riesgo con la temperatura.
A sample composed only of failures cannotestimate the probability of failure.Available evidence could not establish the probability of failure at 31 °F calculated onlyfrom the 7 incident flights, because that subset eliminates the non-damaged class.Sources: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Albert, A., &Anderson, J. A. (1984). On the existence of maximum likelihood estimates in logistic regression models. Biometrika.
If only 7 flights with incidents are selected, each group shows 100% damage by construction. It is not a prediction, but a consequence of the filter. Without negative outcomes, binary logistic regression cannot identify a valid probability at 31 °F.
Una muestra compuesta solo por fallas nopuede estimar la probabilidad de fallar.La evidencia disponible no pudo establecer la probabilidad de falla a 31 °F calculadaúnicamente con los 7 vuelos incidentados, porque ese subconjunto elimina la clase sindaño.Fuentes: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Albert, A., &Anderson, J. A. (1984). On the existence of maximum likelihood estimates in logistic regression models. Biometrika.
Si se seleccionan únicamente 7 vuelos con incidentes, cada grupo presenta 100% de daño por construcción. No es una predicción, sino una consecuencia del filtro. Sin resultados negativos, una regresión logística binaria no identifica una probabilidad válida a 31 °F.
With no flight experience below 53 °F, thedecision required demonstrating thatlaunch would be unsafe.The failure was not merely statistical. The decision structure determined which side hadto prove its case and who could accept an unprecedented extrapolation.Sources: Presidential Commission on the Space Shuttle Challenger Accident. (1986). Report to the President, Volume I, Chapter 5. U.S. GovernmentPrinting Office.
No flight experience existed below 53 °F. Yet the discussion stopped requiring evidence of safety and shifted to demanding conclusive proof of danger. That change converted extreme uncertainty into operating permission.
Sin experiencia bajo 53 °F, ladecisión exigió demostrar quelanzar sería inseguro.La falla no fue únicamente estadística. La estructura de decisión determinó qué ladodebía probar su caso y quién podía aceptar una extrapolación sin precedentes.Fuentes: Presidential Commission on the Space Shuttle Challenger Accident. (1986). Report to the President, Volume I, Chapter 5. U.S. GovernmentPrinting Office.
No existía experiencia de vuelo por debajo de 53 °F. Aun así, la discusión dejó de exigir evidencia de seguridad y pasó a exigir una prueba concluyente de peligro. Ese cambio convirtió una incertidumbre extrema en permiso operativo.
Feynman said engineers and managers estimatedshuttle failure risk 1,000-fold differentlyFailure-with-loss estimates reported by Feynman for technical and managerial staff.Made for Johannes Talero, by HanademiSources: nasa.gov.Frequencies of 1 in 100 and 1 in 100,000 are expressed as 1% and 0%.Unit% probabilityEngineers1%Management0.001%1,000x
Feynman found estimates near 1 in 100 among engineers and 1 in 100,000 among management. That 1,000x gap is not a minor calibration disagreement. It is evidence that information transformed as it rose through the organization.
Sources
Según Feynman, ingenieros y directivos estimaban1.000 veces distinto el riesgo del transbordadorEstimaciones de falla con pérdida reportadas por Feynman para personal técnico y gerencial.Made for Johannes Talero, by HanademiFuentes: nasa.gov.Las frecuencias 1 en 100 y 1 en 100 se expresan como 1% y 0.001%.Unidad% de probabilidadIngenieros1 %Gerencia0,001 %1.000x
Feynman encontró estimaciones cercanas a 1 en 100 entre ingenieros y 1 en 100,000 entre la gerencia. Esa brecha de 1,000 veces no es un desacuerdo menor de calibración. Es evidencia de que la información se transformaba mientras ascendía por la organización.
Fuentes
The decision had human stakesSeven people wearing matching NASA flight suits walk together outdoors.Sources: A Statistical Analysis of The Challenger Accident. rebellionresearch.com.
Seven people in NASA flight suits make the consequences of unresolved risk disagreement impossible to abstract away.
La decisión tenía consecuencias humanasSiete personas con trajes de vuelo de la NASA a juego caminan juntas al aire libre.Fuentes: A Statistical Analysis of The Challenger Accident. rebellionresearch.com.
Siete personas con trajes de vuelo de la NASA hacen imposible abstraer las consecuencias del desacuerdo sobre el riesgo.
Google Flu Trends overestimated flu levelsin 100 of 108 weeksWeeks analyzed and weeks in which Google Flu Trends exceeded CDC reports; evaluation published in 2014.Made for Johannes Talero, by HanademiSources: Ginsberg, J., et al. (2009). Detecting influenza epidemics using search engine query data. Nature, 457, 1012-1014.; Lazer, D.,Kennedy, R., King, G., & Vespignani, A. (2014). The parable of Google Flu: Traps in big data analysis. Science.UnitweeksAnalyzed108Overestimated100Not overestimated8
Google Flu Trends began with a median correlation of 0.97 against weekly CDC surveillance. In 2013 it overestimated the peak by 140% and remained above official reports for 100 of 108 weeks. The digital signal worked better as a complement to epidemiological surveillance, not as a replacement.
Google Flu Trends sobrestimó los niveles degripe en 100 de 108 semanasSemanas analizadas y semanas en que Google Flu Trends superó los reportes de los CDC; evaluaciónpublicada en 2014.Made for Johannes Talero, by HanademiFuentes: Ginsberg, J., et al. (2009). Detecting influenza epidemics using search engine query data. Nature, 457, 1012-1014.; Lazer, D.,Kennedy, R., King, G., & Vespignani, A. (2014). The parable of Google Flu: Traps in big data analysis. Science.UnidadsemanasAnalizadas108Sobrestimadas100No sobrestimadas8
Google Flu Trends comenzó con una correlación media de 0.97 frente a la vigilancia semanal de los CDC. En 2013 sobrestimó 140% el pico y quedó por encima de los reportes oficiales durante 100 de 108 semanas. La señal digital funcionaba mejor como complemento de la vigilancia epidemiológica, no como reemplazo.
Zillow’s home purchases rose from 1,856to 9,680 in 2021Homes purchased by Zillow Offers during the first 3 quarters of 2021.Made for Johannes Talero, by HanademiSources: Zillow Group, Inc. (2021). Quarterly shareholder letters for the first, second and third quarters of 2021. Zillow Group.Unithomes05K10K2021-T12021-T22021-T3In 2021-Q3, Zillow purchased9,680 homes.1,856
Zillow Offers accelerated its purchases during 2021. It moved from 1,856 homes in the first quarter to 9,680 in the third. Each pricing error stopped being a wrong forecast and became real inventory on the balance sheet.
Las compras de viviendas de Zillowsubieron de 1.856 a 9.680 en 2021Viviendas compradas por Zillow Offers durante los 3 primeros trimestres de 2021.Made for Johannes Talero, by HanademiFuentes: Zillow Group, Inc. (2021). Quarterly shareholder letters for the first, second and third quarters of 2021. Zillow Group.Unidadviviendas05K10K2021-T12021-T22021-T3En 2021-T3, Zillow compró9,680 viviendas.1.856
Zillow Offers aceleró sus compras durante 2021. Pasó de 1,856 viviendas en el primer trimestre a 9,680 en el tercero. Cada error de precio dejó de ser una predicción equivocada y se convirtió en inventario real sobre el balance.
Zillow recorded a US$304 million chargefor homes that had lost valueInventory impairment recorded in 2021-Q3 and projected range for 2021-Q4, millions of USD.Made for Johannes Talero, by HanademiSources: Zillow Group, Inc. (2021). Third quarter 2021 shareholder letter. Zillow Group.Unitmillions of USD$3042021-Q3 recorded$2652021-Q4 maximum$2402021-Q4 minimum
The cost appeared when price estimates met actual buyers. Zillow recorded USD 304 million in impairments during 2021-Q3. It also anticipated between USD 240 and 265 million for 2021-Q4, which is why those 2 values are presented as forecasts rather than observed losses.
Zillow registró un cargo de US$304 millones porviviendas que habían perdido valorDeterioro de inventario registrado en 2021-T3 y rango previsto para 2021-T4, millones de USD.Made for Johannes Talero, by HanademiFuentes: Zillow Group, Inc. (2021). Third quarter 2021 shareholder letter. Zillow Group.Unidadmillones de USD$3042021-T3 registrado$2652021-T4 máximo$2402021-T4 mínimo
El costo apareció cuando las estimaciones de precios encontraron compradores reales. Zillow registró USD 304 millones en deterioros durante 2021-T3. También anticipó entre USD 240 y 265 millones para 2021-T4, por eso esos 2 valores se presentan como previsiones y no como pérdidas observadas.
The model failed on nodashboard: it affected inventory,strategy, and employment.The 25% is the approximate reduction announced by Zillow, not a subsequentmeasurement of executed layoffs.Sources: Zillow Group, Inc. (2021, November 2). Zillow Group reports third-quarter 2021 financial results and announces wind-down of Zillow Offers.Zillow Group.
Zillow closed Offers after acknowledging it could not predict prices with the precision needed to operate at that scale. The company announced an approximate 25% reduction in its workforce. The predictive error had acquired irreversible operational consequences.
El modelo no falló en un tablero: afectóinventario, estrategia y empleo.El 25% es la reducción aproximada anunciada por Zillow, no una medición posterior dedespidos ejecutados.Fuentes: Zillow Group, Inc. (2021, November 2). Zillow Group reports third-quarter 2021 financial results and announces wind-down of Zillow Offers.Zillow Group.
Zillow cerró Offers después de reconocer que no podía prever los precios con la precisión necesaria para operar a esa escala. La empresa anunció una reducción aproximada de 25% de su plantilla. El error predictivo había adquirido consecuencias operativas irreversibles.
A study found face classifiers mislabeledgender in up to 34.7% of imagesGender classification error for dark-skinned women across 3 commercial systems evaluated by GenderShades.Made for Johannes Talero, by HanademiSources: Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercial gender classification.Proceedings of Machine Learning Research, 81, 77-91.Unit% error0%20%40%34.7%IBM20.8%Microsoft34.5%Face++Maximum in light-skinned men
Gender Shades disaggregated performance that an average could obscure. Errors for dark-skinned women reached 34.7%, 20.8%, and 34.5%. In light-skinned men, none of the 3 systems exceeded 0.8% error.
Un estudio halló que clasificadores faciales erraron elgénero hasta en 34,7% de imágenesError de clasificación de género para mujeres de piel oscura en 3 sistemas comerciales evaluados porGender Shades.Made for Johannes Talero, by HanademiFuentes: Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercial gender classification.Proceedings of Machine Learning Research, 81, 77-91.Unidad% de error0 %20 %40 %34,7 %IBM20,8 %Microsoft34,5 %Face++Máximo en hombres de piel clara
Gender Shades desagregó el rendimiento que una media podía ocultar. Los errores para mujeres de piel oscura llegaron a 34.7%, 20.8% y 34.5%. En hombres de piel clara, ninguno de los 3 sistemas superó 0.8% de error.
Gender-labeling errors for dark-skinned womenwere 26–43 times light-skinned men’sDivide each dark-skinned-women error rate by the maximum reported light-skinned-men error rate of0.8%.Made for Johannes Talero, by HanademiSources: Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercialgender classification. Proceedings of Machine Learning Research, 81, 77-91.; Results ‹ Gender Shades.UnitMinimum error multiple versus light-skinned menIBM43.4Microsoft26Face++43.1
Using the most conservative denominator still produces error gaps above 26 times for all three commercial classifiers.
Mujeres de piel oscura tuvieron 26–43 veces máserrores de género que hombres de piel claraDividir cada tasa de error en mujeres de piel oscura entre la tasa máxima reportada de 0.8% en hombres depiel clara.Made for Johannes Talero, by HanademiFuentes: Buolamwini, J., & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercialgender classification. Proceedings of Machine Learning Research, 81, 77-91.; Results ‹ Gender Shades.UnidadMúltiplo mínimo del error frente a hombres de piel claraIBM43,4Microsoft26Face++43,1
Incluso con el denominador más conservador, los tres clasificadores comerciales muestran brechas de error superiores a 26 veces.
The sepsis model detected only 33% of casesSensitivity, specificity, and positive predictive value in an external hospital validation published in 2021.Made for Johannes Talero, by HanademiSources: Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMAInternal Medicine.Unit%Sensitivity33%Specificity83%Positive predictive value12%
An aggregated figure could make the model appear competent. But sensitivity was 33% and positive predictive value barely 12%. In a clinical alert, performance must be read from the harm each error type produces.
El modelo de sepsis detectó solo 33% de loscasosSensibilidad, especificidad y valor predictivo positivo en una validación hospitalaria externa publicada en2021.Made for Johannes Talero, by HanademiFuentes: Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMAInternal Medicine.Unidad%Sensibilidad33 %Especificidad83 %Valor predictivo positivo12 %
Una cifra agregada podía hacer que el modelo pareciera competente. Pero la sensibilidad fue 33% y el valor predictivo positivo apenas 12%. En una alerta clínica, el rendimiento debe leerse desde el daño que produce cada tipo de error.
The sepsis model missed 67% of caseswhile 88% of its alerts were falseMade for Johannes Talero, by HanademiSources: Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction modelin hospitalized patients. JAMA Internal Medicine.UnitPercentage within each evaluation groupPatients who developed sepsis33% detected67% missedPatients without sepsis83% correctly remained quiet17% false-positive rateAlerts issued12% true alerts88% false alerts
Strong specificity did not prevent failure on both operational sides: most sepsis cases were missed and most issued alerts were false.
El modelo de sepsis omitió 67% de los casosy 88% de sus alertas fueron falsasMade for Johannes Talero, by HanademiFuentes: Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction modelin hospitalized patients. JAMA Internal Medicine.UnidadPorcentaje dentro de cada grupo de evaluaciónPacientes que desarrollaron sepsis33% detectados67% omitidosPacientes sin sepsis83% sin alerta correctamente17% de falsos positivosAlertas emitidas12% de alertas verdaderas88% de alertas falsas
La alta especificidad no evitó fallas en ambos frentes operativos: se omitió la mayoría de los casos y la mayoría de las alertas fueron falsas.
One hospital sepsis model gave no timelyalert for 67% of casesSepsis cases detected and cases without timely alert in the evaluated hospital implementation.Made for Johannes Talero, by HanademiSources: Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients.JAMA Internal Medicine.Unit% of casesCases detectedWithout timely alert33%67%
A sensitivity of 33% has a direct translation. For every cohort of patients who developed sepsis, 67% did not receive a timely alert. The critical metric was not overall average but how many cases went unwarned.
Un modelo hospitalario de sepsis no alertó atiempo en 67% de los casosCasos de sepsis detectados y casos sin alerta oportuna en la implementación hospitalaria evaluada.Made for Johannes Talero, by HanademiFuentes: Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients.JAMA Internal Medicine.Unidad% de casosCasos detectadosSin alerta oportuna33 %67 %
La sensibilidad de 33% tiene una traducción directa. Por cada grupo de pacientes que desarrolló sepsis, 67% no recibió una alerta oportuna. La métrica decisiva no era el promedio general, sino cuántos casos quedaban sin advertencia.
With newly collected images, accuracy at labelingobjects fell by up to 14.3 percentage pointsReported accuracy drop in an ImageNet replica, percentage points; does not compare complete samplesagainst incident-only samples.Made for Johannes Talero, by HanademiSources: Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet? Proceedings ofMachine Learning Research, 97, 5389-5400.Unitpercentage pointsMinimum reported11.7Midpoint13Maximum reported14.3
Known models performed worse when ImageNet was replicated with new data. The reported drop was between 11.7 and 14.3 percentage points. The result demonstrates degradation outside the original test, but does not answer the exact comparison between complete data and incident-only samples.
Con imágenes recién recopiladas, la exactitud aletiquetar objetos cayó hasta 14,3 puntos porcentualesCaída reportada de exactitud en una réplica de ImageNet, puntos porcentuales; no compara muestrascompletas contra muestras solo de incidentes.Made for Johannes Talero, by HanademiFuentes: Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet? Proceedingsof Machine Learning Research, 97, 5389-5400.Unidadpuntos porcentualesMínimo reportado11,7Punto medio13Máximo reportado14,3
Los modelos conocidos rindieron peor cuando ImageNet se replicó con datos nuevos. La caída reportada estuvo entre 11.7 y 14.3 puntos porcentuales. El resultado demuestra degradación fuera de la prueba original, pero no responde la comparación exacta entre datos completos y muestras solo de incidentes.
Available evidence could not establish 3requested critical measures.Available evidence could not establish: the failure probability at 31 °F calculated fromonly the 7 incident flights. Available evidence could not establish: the number ofdocumented failures in the AI Incident Database attributed to out-of-distributionextrapolation. Available evidence could not establish: the productive degradationdifference between models evaluated on complete data and models evaluated on incidentsalone.Sources: Responsible AI Collaborative. (2026). AI Incident Database. Responsible AI Collaborative.; Responsible AI Collaborative. (2026). About the AIIncident Database. Responsible AI Collaborative.; National Institute of Standards and Technology. (2023). Artificial Intelligence Risk ManagementFramework 1.0. U.S. Department of Commerce.
Available evidence could not establish a valid probability at 31 °F using only the 7 incidents. It could not count AI failures caused specifically by out-of-distribution extrapolation. It found no consolidated productive comparison between models evaluated on complete data and models evaluated on incidents alone.
La evidencia disponible no pudo establecer3 medidas críticas solicitadas.La evidencia disponible no pudo establecer: la probabilidad de falla a 31 °F calculada solocon los 7 vuelos incidentados. La evidencia disponible no pudo establecer: el número defallas documentadas en AI Incident Database atribuidas a extrapolación fuera dedistribución. La evidencia disponible no pudo establecer: la diferencia de degradaciónproductiva entre modelos evaluados con datos completos y modelos evaluados solo conincidentes.Fuentes: Responsible AI Collaborative. (2026). AI Incident Database. Responsible AI Collaborative.; Responsible AI Collaborative. (2026). About the AIIncident Database. Responsible AI Collaborative.; National Institute of Standards and Technology. (2023). Artificial Intelligence Risk ManagementFramework 1.0. U.S. Department of Commerce.
La evidencia disponible no pudo establecer una probabilidad válida a 31 °F usando solo los 7 incidentes. Tampoco pudo contar los incidentes de IA causados específicamente por extrapolación fuera de distribución. Finalmente, no encontró una comparación productiva consolidada entre modelos evaluados con datos completos y modelos evaluados solo con incidentes.
NIST requires 4 functions: govern, map,measure, and manage risk.The practical protocol is to preserve denominators, measure subgroups, flagextrapolations, define operational limits, and assign authority to halt deployment.Sources: National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework 1.0. U.S. Department ofCommerce.
NIST converts the Challenger lessons into 4 operational verbs: govern, map, measure, and manage. Order matters because measuring without authority does not stop a deployment. Each team needs explicit boundaries and a person with the power to say no.
NIST exige 4 funciones: gobernar, mapear,medir y gestionar el riesgo.El protocolo práctico es conservar denominadores, medir subgrupos, marcarextrapolaciones, definir límites operativos y asignar autoridad para detener eldespliegue.Fuentes: National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework 1.0. U.S. Department ofCommerce.
NIST convierte las lecciones del Challenger en 4 verbos operativos: gobernar, mapear, medir y gestionar. El orden importa porque medir sin autoridad no detiene un despliegue. Cada equipo necesita límites explícitos y una persona con poder para decir no.
In summaryThe 5 findings that must survive the presentation.Made for Johannes Talero, by HanademiSources: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Lazer, D., Kennedy,R., King, G., & Vespignani, A. (2014). The parable of Google Flu: Traps in big data analysis. Science.; Wong, A., et al. (2021). External validation of awidely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine.31°F was 22°F outside the range.Excluding 16 flights erased the decisive denominator.Google overestimated in 100 of 108 weeks.Prediction error ended up cutting 25% of headcount.The model left 67% of cases without alert.
Challenger is not a story about an ugly graph or a single wrong calculation. It is a story about excluding denominators, extrapolating without limits, and shifting the burden of proof to those warning of danger. Modern cases change industries but preserve the same structure.
En resumenLas 5 conclusiones que deben sobrevivir a la presentación.Made for Johannes Talero, by HanademiFuentes: UCI Machine Learning Repository. (1993). Challenger USA Space Shuttle O-Ring dataset. University of California, Irvine.; Lazer, D., Kennedy,R., King, G., & Vespignani, A. (2014). The parable of Google Flu: Traps in big data analysis. Science.; Wong, A., et al. (2021). External validation of awidely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine.31 °F quedó 22 °F fuera del rango.Excluir 16 vuelos borró el denominador decisivo.Google sobrestimó 100 de 108 semanas.El error predictivo terminó recortando 25% de la plantilla.El modelo dejó 67% de casos sin alerta.
Challenger no es una historia sobre una gráfica fea ni sobre un único cálculo errado. Es una historia sobre excluir denominadores, extrapolar sin límites y trasladar la carga de la prueba a quienes advertían del peligro. Los casos modernos cambian de industria, pero conservan la misma estructura.

The research behind this deck

A forensic reconstruction of how incomplete data, extreme extrapolation, and organizational pressure authorized Challenger, and how those same errors resurface in AI systems deployed at scale.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

Una reconstrucción forense de cómo datos incompletos, extrapolación extrema y presión organizacional autorizaron el Challenger, y cómo esos mismos errores reaparecen en sistemas de IA desplegados a escala.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English

Related researchInvestigación relacionada