Hanademi

The AI engineer roadmap that survives production

26 slides · 22 min · 1 minute ago language
ENES
theme
LightDark
view
DeckTableTalk
brand
HanademiPlatzi
The AI engineer roadmap thatsurvives productionMade for Oscar Ruiz, by Hanademi
El roadmap de AI Engineer queresiste producciónMade for Oscar Ruiz, by Hanademi
Practice explains just 1% of performance inprofessionsShare of performance variance attributed to deliberate practice in a 2014 meta-analysis.Made for Oscar Ruiz, by HanademiSources: Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports,education, and professions. Psychological Science.Unit% of variance26%Games21%Music18%Sports4%Education1%Professions
A fixed number of hours cannot promise professional competence. The meta-analysis attributed only 1% of performance variance in professions to deliberate practice. A roadmap therefore needs evidence of completed, tested and maintained work.
La práctica explica solo 1% del rendimientoprofesionalPorcentaje de la varianza del rendimiento atribuido a la práctica deliberada en un metaanálisis de 2014.Made for Oscar Ruiz, by HanademiFuentes: Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports,education, and professions. Psychological Science.Unidad% de la varianza26 %Juegos21 %Música18 %Deportes4 %Educación1 %Profesiones
Un número fijo de horas no puede prometer competencia profesional. El metaanálisis atribuyó solo 1% de la varianza del rendimiento profesional a la práctica deliberada. Por eso, un roadmap necesita pruebas de trabajo terminado, probado y mantenido.
The six terms worth knowingMade for Oscar Ruiz, by HanademiLLMA language model trained to generate and interpret text.RAGA system that retrieves evidence before generating an answer.LoRAA method that adjusts small adapters instead of every parameter.QLoRALoRA with lower-precision model weights to save memory.vLLMSoftware designed to serve language models efficiently.MCPAn open protocol for connecting AI applications to tools and data.
The roadmap crosses several layers of the AI stack. These terms separate retrieval, adaptation, serving and interoperability. Professional judgment begins with knowing which problem each layer solves.
Los seis términos que conviene conocerMade for Oscar Ruiz, by HanademiLLMUn modelo de lenguaje entrenado para generar e interpretar texto.RAGUn sistema que recupera evidencia antes de generar una respuesta.LoRAUn método que ajusta pequeños adaptadores en vez de todos los parámetros.QLoRALoRA con pesos de menor precisión para ahorrar memoria.vLLMSoftware diseñado para servir modelos de lenguaje con eficiencia.MCPUn protocolo abierto para conectar aplicaciones de IA con herramientas y datos.
El roadmap cruza varias capas del stack de IA. Estos términos separan recuperación, adaptación, servicio e interoperabilidad. El criterio profesional empieza por saber qué problema resuelve cada capa.
The same 1,000 hours can mean 25 or 100weeksWeeks required to accumulate 1,000 hours at 3 weekly schedules, before holidays or interruptions.Made for Oscar Ruiz, by HanademiSources: edworkingpapers.com.; cmc.edu.Planning calculation: 1,000 divided by weekly hours.UnitWeeks10 hours/week10022.5 hours/week44.440 hours/week25
A 1,000-hour target becomes 100 weeks at 10 weekly hours and 25 weeks at 40. Time still has consequences: cutting 2 school weeks harmed performance in Madrid, while a 150-hour professional rule reduced early CPA candidates. Duration matters, but more required time is not automatically better.
Las mismas 1.000 horas pueden significar25 o 100 semanasSemanas necesarias para acumular 1.000 horas con 3 ritmos semanales, antes de vacaciones ointerrupciones.Made for Oscar Ruiz, by HanademiFuentes: edworkingpapers.com.; cmc.edu.Cálculo de planificación: 1.000 dividido por las horas semanales.UnidadSemanas10 horas/semana10022,5 horas/semana44,440 horas/semana25
Una meta de 1.000 horas se convierte en 100 semanas con 10 horas semanales y en 25 con 40. El tiempo sí tiene consecuencias: recortar 2 semanas escolares perjudicó el rendimiento en Madrid, mientras una regla profesional de 150 horas redujo los primeros candidatos a contador. La duración importa, pero exigir más tiempo no es automáticamente mejor.
AI and big data lead the fastest-growingskillsWorld Economic Forum ranking of the 5 fastest-growing skill groups through 2030; lower rank is better.Made for Oscar Ruiz, by HanademiSources: World Economic Forum. (2025). The Future of Jobs Report 2025. World Economic Forum.UnitRank5Resilience and flexibility4Creative thinking3Technological literacy2Networks and cybersecurity1AI and big data
AI and big data ranked first among the fastest-growing skills through 2030. Networks and cybersecurity ranked second, while technological literacy ranked third. The market is asking for a stack of connected abilities, not prompt writing alone.
IA y big data lideran las habilidades demayor crecimientoRanking del Foro Económico Mundial de los 5 grupos de habilidades con mayor crecimiento hasta 2030; unrango menor es mejor.Made for Oscar Ruiz, by HanademiFuentes: World Economic Forum. (2025). The Future of Jobs Report 2025. World Economic Forum.UnidadPosición5Resiliencia y flexibilidad4Pensamiento creativo3Alfabetización tecnológica2Redes y ciberseguridad1IA y big data
IA y big data ocuparon el primer lugar entre las habilidades de mayor crecimiento hasta 2030. Redes y ciberseguridad quedaron segundas y alfabetización tecnológica, tercera. El mercado pide un conjunto de capacidades conectadas, no solo escribir prompts.
AI demand rises behind a narrow juniorgateGlobal employer expectations through 2030 and Indian technology postings tracked from May to July 2026.Made for Oscar Ruiz, by HanademiSources: World Economic Forum. (2025). The Future of Jobs Report 2025. World Economic Forum.; prepnplaced.com.UnitPercent and postingsEmployer expectations86%77%41%AI transformsbusinessPlan reskillingExpect staff cutsIndian technology postings15,429908All postingsEntry or junior
Among surveyed employers, 86% expected AI to transform their business and 77% planned reskilling. Entry access was far tighter in one 2026 Indian dataset, where only 908 of 15,429 postings were entry or junior roles. The portfolio must reduce hiring risk.
La demanda de IA crece tras una puertajunior estrechaExpectativas globales de empleadores hasta 2030 y vacantes tecnológicas de India monitoreadas entremayo y julio de 2026.Made for Oscar Ruiz, by HanademiFuentes: World Economic Forum. (2025). The Future of Jobs Report 2025. World Economic Forum.; prepnplaced.com.UnidadPorcentaje y vacantesExpectativas empresariales86 %77 %41 %IA transforma elnegocioPlanean recualificarPrevén recortesVacantes tecnológicas en India15.429908Todas las vacantesEntrada o junior
Entre los empleadores encuestados, 86% esperaba que la IA transformara su negocio y 77% planeaba recualificación. El acceso inicial fue mucho más estrecho en un conjunto indio de 2026, donde solo 908 de 15.429 vacantes eran de entrada o junior. El portafolio debe reducir el riesgo de contratación.
Python leads AI jobs, but SQL still mattersProgramming languages appearing in 1,969 tracked AI job postings, percent of postings.Made for Oscar Ruiz, by HanademiSources: theaimarketpulse.com.; GitHub. (2024). Octoverse 2024: The state of open source and rise of AI. GitHub.Unit% of AI postingsPython65%SQL38%JavaScript/TypeScript24%Java12%Rust6%C++5%
Python appeared in 65% of tracked AI postings, but SQL still appeared in 38%. JavaScript or TypeScript appeared in 24%. GitHub also reported that Python became its most-used language in 2024.
Python lidera el empleo en IA, pero SQLsigue importandoLenguajes presentes en 1.969 vacantes de IA monitoreadas, porcentaje de vacantes.Made for Oscar Ruiz, by HanademiFuentes: theaimarketpulse.com.; GitHub. (2024). Octoverse 2024: The state of open source and rise of AI. GitHub.Unidad% de vacantes de IAPython65 %SQL38 %JavaScript/TypeScript24 %Java12 %Rust6 %C++5 %
Python apareció en 65% de las vacantes de IA monitoreadas, pero SQL todavía apareció en 38%. JavaScript o TypeScript apareció en 24%. GitHub también informó que Python se convirtió en su lenguaje más utilizado en 2024.
PostgreSQL leads, but the storage stackstays mixedDatabase usage in Stack Overflow's 2024 Developer Survey, approximate share of respondents.Made for Oscar Ruiz, by HanademiSources: Stack Overflow. (2024). 2024 Developer Survey: Databases. Stack Overflow.Unit% of respondents49%PostgreSQL41%MySQL33%SQLite25%SQL Server20%Redis20%MongoDB
PostgreSQL had the largest reported share at about 49%, but no database erased the others. SQLite and Redis solve different operational problems from PostgreSQL. Professional judgment means matching persistence, portability and latency to the workload.
PostgreSQL lidera, pero el stack de datossigue mezcladoUso de bases de datos en la encuesta de desarrolladores de Stack Overflow de 2024, proporciónaproximada de participantes.Made for Oscar Ruiz, by HanademiFuentes: Stack Overflow. (2024). 2024 Developer Survey: Databases. Stack Overflow.Unidad% de participantes49 %PostgreSQL41 %MySQL33 %SQLite25 %SQL Server20 %Redis20 %MongoDB
PostgreSQL tuvo la mayor proporción informada, cerca de 49%, pero ninguna base eliminó a las demás. SQLite y Redis resuelven problemas operativos distintos de PostgreSQL. El criterio profesional consiste en adaptar persistencia, portabilidad y latencia a la carga.
The FastAPI refactor improved throughputand latency togetherBefore and after one reported application refactor; throughput and p95 latency use separate panels.Made for Oscar Ruiz, by HanademiSources: dev.to.UnitRequests per second and millisecondsThroughput1801,300BeforeAfterp95 latency4,000200Before, overAfter, under
One reported FastAPI refactor raised throughput from about 180 to 1,300 requests per second. It also reduced p95 latency from more than 4 seconds to less than 200 ms. The useful lesson is end-to-end profiling, not a universal framework ranking.
Sources
La refactorización de FastAPI mejorórendimiento y latenciaAntes y después de una refactorización reportada; el rendimiento y la latencia p95 usan paneles separados.Made for Oscar Ruiz, by HanademiFuentes: dev.to.UnidadSolicitudes por segundo y milisegundosRendimiento1801.300AntesDespuésLatencia p954.000200Antes, más deDespués, menos de
Una refactorización reportada de FastAPI elevó el rendimiento de unas 180 a 1.300 solicitudes por segundo. También redujo la latencia p95 de más de 4 segundos a menos de 200 ms. La lección útil es perfilar todo el sistema, no crear un ranking universal de frameworks.
Fuentes
Argentina's fixed broadband base grewelevenfoldFixed broadband subscriptions per 100 people in Argentina, 2005 to 2024.Made for Oscar Ruiz, by HanademiSources: World Bank. (n.d.). Fixed broadband subscriptions (per 100 people) [IT.NET.BBND.P2]. World Bank OpenData.UnitSubscriptions per 100 people010202005200820112014201720202023Argentina reached 26.1subscriptions per 100 people in2024.2.4
Argentina moved from 2.36 fixed broadband subscriptions per 100 people in 2005 to 26.1 in 2024. That is roughly an elevenfold expansion. Access supports remote learning and deployment, although it does not determine professional competence.
La banda ancha fija de Argentina crecióonce vecesSuscripciones de banda ancha fija por cada 100 personas en Argentina, 2005 a 2024.Made for Oscar Ruiz, by HanademiFuentes: World Bank. (n.d.). Fixed broadband subscriptions (per 100 people) [IT.NET.BBND.P2]. World Bank OpenData.UnidadSuscripciones por cada 100 personas010202005200820112014201720202023Argentina alcanzó 26,1suscripciones por cada 100personas en 2024.2,4
Argentina pasó de 2,36 suscripciones de banda ancha fija por cada 100 personas en 2005 a 26,1 en 2024. Es una expansión de aproximadamente once veces. El acceso facilita el aprendizaje y despliegue remotos, aunque no determina la competencia profesional.
Internet use in Argentina climbed from 7%to nearly 90%Individuals using the Internet as a share of Argentina's population, 2000 to 2024.Made for Oscar Ruiz, by HanademiSources: World Bank. (n.d.). Individuals using the Internet (% of population) [IT.NET.USER.ZS]. World Bank Open Data.Unit% of population0%50%100%2000200420082012201620202024Internet use reached 89.7% in2024.
Internet use in Argentina rose from about 7% in 2000 to 89.7% in 2024. Access to documentation, courses and cloud tools is far broader than it was one generation ago. The scarce asset is increasingly the ability to turn that access into working systems.
El uso de Internet en Argentina subió de 7%a casi 90%Personas que usan Internet como proporción de la población de Argentina, 2000 a 2024.Made for Oscar Ruiz, by HanademiFuentes: World Bank. (n.d.). Individuals using the Internet (% of population) [IT.NET.USER.ZS]. World Bank Open Data.Unidad% de la población0 %50 %100 %2000200420082012201620202024El uso de Internet alcanzó89,7% en 2024.
El uso de Internet en Argentina pasó de cerca de 7% en 2000 a 89,7% en 2024. El acceso a documentación, cursos y herramientas cloud es mucho más amplio que hace una generación. El recurso escaso es cada vez más la capacidad de convertir ese acceso en sistemas funcionales.
Argentina's secure-server density rosemore than 200-foldSecure Internet servers per 1 million people in Argentina, 2010 to 2024. Log scale: every gridline is ×10.Made for Oscar Ruiz, by HanademiSources: World Bank. (n.d.). Secure Internet servers (per 1 million people) [IT.NET.SECR.P6]. World Bank Open Data.UnitServers per 1 million people101001,00010,00020102013201620192022The 2024 level reached 5,453servers per million people.24.9
Argentina moved from about 25 secure Internet servers per million people in 2010 to 5,453 in 2024. The infrastructure supporting online applications expanded dramatically. More infrastructure also creates more systems that require secure configuration, monitoring and maintenance.
La densidad de servidores seguros deArgentina creció más de 200 vecesServidores seguros de Internet por cada millón de personas en Argentina, 2010 a 2024. Escalalogarítmica: cada línea es ×10.Made for Oscar Ruiz, by HanademiFuentes: World Bank. (n.d.). Secure Internet servers (per 1 million people) [IT.NET.SECR.P6]. World Bank OpenData.UnidadServidores por millón de personas101001.00010.00020102013201620192022El nivel de 2024 alcanzó 5.453servidores por millón depersonas.24,9
Argentina pasó de unos 25 servidores seguros de Internet por millón de personas en 2010 a 5.453 en 2024. La infraestructura que sostiene aplicaciones en línea creció de forma drástica. Más infraestructura también crea más sistemas que requieren configuración segura, monitoreo y mantenimiento.
Linux fluency transfers directly into thecompute layer beneath AI.Shell, processes, files, permissions and resource inspection are highly transferableproduction skills.Sources: TOP500. (2024). Operating system family statistics. TOP500.; linuxcareers.com.
Linux reached 100% representation in the TOP500 in 2017 and retained it through the cited 2024 endpoint. A separate 2026 analysis found AI mentioned in 18% of 1,675 distinct Linux job postings after duplicate removal. Learn enough Linux to diagnose systems without turning the roadmap into systems administration.
La soltura con Linux se transfieredirectamente a la capa de cómputo de la IA.La terminal, los procesos, los archivos, los permisos y la inspección de recursos sonhabilidades productivas muy transferibles.Fuentes: TOP500. (2024). Operating system family statistics. TOP500.; linuxcareers.com.
Linux alcanzó 100% de representación en el TOP500 en 2017 y la mantuvo hasta el punto citado de 2024. Un análisis separado de 2026 encontró menciones de IA en 18% de 1.675 vacantes Linux distintas tras eliminar duplicados. Aprende suficiente Linux para diagnosticar sistemas sin convertir el roadmap en administración de sistemas.
ICT goods lost share of Argentina's importbasketICT goods as a share of Argentina's total goods imports, 2000 to 2024.Made for Oscar Ruiz, by HanademiSources: World Bank. (n.d.). ICT goods imports (% total goods imports) [TM.VAL.ICTG.ZS.UN]. World Bank Open Data.Unit% of goods imports0%5%10%2000200420082012201620202024ICT goods represented 13.79% in2000.The share fell to 6.04% in 2024.
ICT goods represented 13.79% of Argentina's goods imports in 2000 and 6.04% in 2024. The series does not measure hardware availability directly, but it shows that technology goods occupy a smaller import share than before. Efficient inference and portable deployment remain practical skills.
Los bienes TIC perdieron peso en lasimportaciones argentinasBienes TIC como proporción de las importaciones totales de bienes de Argentina, 2000 a 2024.Made for Oscar Ruiz, by HanademiFuentes: World Bank. (n.d.). ICT goods imports (% total goods imports) [TM.VAL.ICTG.ZS.UN]. World Bank Open Data.Unidad% de importaciones de bienes0 %5 %10 %2000200420082012201620202024Los bienes TIC representaron13,79% en 2000.La proporción cayó a 6,04% en2024.
Los bienes TIC representaron 13,79% de las importaciones argentinas de bienes en 2000 y 6,04% en 2024. La serie no mide directamente la disponibilidad de hardware, pero indica que los bienes tecnológicos ocupan una proporción menor que antes. La inferencia eficiente y el despliegue portátil siguen siendo habilidades prácticas.
India's service exports remain far moredigitalICT services as a share of service exports, 2005 to 2025; Spain begins in 2013.Made for Oscar Ruiz, by HanademiSources: World Bank. (n.d.). ICT service exports (% of service exports, BoP) [BX.GSR.CCIS.ZS]. World Bank Open Data.Unit% of service exports0%20%40%2005200820112014201720202023ArgentinaChileColombiaSpainIndiaMexicoIndia ended 2025 at 48.6%.Argentina peaked at 24.6% in2021.
ICT services represented 48.6% of India's service exports in 2025. Argentina finished at 15.7%, Colombia at 11.8% and Spain at 10.4%. The comparison shows how differently economies translate technical capability into exported services.
Las exportaciones de servicios de Indiasiguen siendo mucho más digitalesServicios TIC como proporción de las exportaciones de servicios, 2005 a 2025; España comienza en2013.Made for Oscar Ruiz, by HanademiFuentes: World Bank. (n.d.). ICT service exports (% of service exports, BoP) [BX.GSR.CCIS.ZS]. World Bank OpenData.Unidad% de exportaciones de servicios0 %20 %40 %2005200820112014201720202023ArgentinaChileColombiaEspañaIndiaMéxicoIndia terminó 2025 en 48,6%.Argentina alcanzó un pico de 24,6%en 2021.
Los servicios TIC representaron 48,6% de las exportaciones de servicios de India en 2025. Argentina terminó en 15,7%, Colombia en 11,8% y España en 10,4%. La comparación muestra cómo las economías convierten de forma distinta la capacidad técnica en servicios exportados.
Colombia pulled far ahead in mobilesubscriptionsMobile cellular subscriptions per 100 people across 6 countries, 2000 to 2024.Made for Oscar Ruiz, by HanademiSources: World Bank. (n.d.). Mobile cellular subscriptions (per 100 people) [IT.CEL.SETS.P2]. World Bank Open Data.UnitSubscriptions per 100 people0501001502000200420082012201620202024ArgentinaChileColombiaSpainMexicoUnitedStates
Every country in the comparison ended above 100 mobile subscriptions per 100 people. Colombia reached 174.1 in 2024, far above the United States at 113.2. Multimodal products can reach broad device populations, but access does not guarantee equal performance across languages or conditions.
Colombia se adelantó ampliamente ensuscripciones móvilesSuscripciones móviles por cada 100 personas en 6 países, 2000 a 2024.Made for Oscar Ruiz, by HanademiFuentes: World Bank. (n.d.). Mobile cellular subscriptions (per 100 people) [IT.CEL.SETS.P2]. World Bank OpenData.UnidadSuscripciones por cada 100 personas0501001502000200420082012201620202024ArgentinaChileColombiaEspañaMéxicoEstadosUnidos
Todos los países de la comparación terminaron por encima de 100 suscripciones móviles por cada 100 personas. Colombia alcanzó 174,1 en 2024, muy por encima de Estados Unidos con 113,2. Los productos multimodales pueden llegar a poblaciones amplias, pero el acceso no garantiza igual rendimiento entre idiomas o condiciones.
Memory engineering can transform modeleconomicsMaximum vLLM throughput gains against 2 serving baselines in tested workloads.Made for Oscar Ruiz, by HanademiSources: Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. Proceedings of the 29th ACM Symposium on Operating SystemsPrinciples.; Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural Information…Sources use different measurement bases; read the comparison directionally, not as one exact scale.UnitThroughput multipleHugging Face Transformers24FasterTransformer3.56.9x
vLLM reported maximum throughput gains of 24 times over Hugging Face Transformers and 3.5 times over FasterTransformer. FlashAttention separately reported training up to 3 times faster while reducing attention-memory growth from quadratic to linear. Both results show why systems knowledge belongs in an AI-engineering roadmap.
La ingeniería de memoria puedetransformar la economía del modeloGanancias máximas de rendimiento de vLLM frente a 2 baselines de servicio en cargas evaluadas.Made for Oscar Ruiz, by HanademiFuentes: Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. Proceedings of the 29th ACM Symposium on OperatingSystems Principles.; Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural…Las fuentes usan bases de medición distintas; lea la comparación como tendencia, no como una escala exacta.UnidadMúltiplo de rendimientoHugging Face Transformers24FasterTransformer3,56,9x
vLLM informó ganancias máximas de rendimiento de 24 veces frente a Hugging Face Transformers y 3,5 veces frente a FasterTransformer. FlashAttention informó por separado entrenamiento hasta 3 veces más rápido y redujo de cuadrático a lineal el crecimiento de memoria de la atención. Ambos resultados muestran por qué los conocimientos de sistemas pertenecen al roadmap.
QLoRA reached 65B parameters on one GPUModel sizes fine-tuned in the QLoRA work; the 65B model used one 48 GB GPU with 4-bit base weights.Made for Oscar Ruiz, by HanademiSources: Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs.Advances in Neural Information Processing Systems, 36.UnitBillion parameters7B713B1333B3365B65
QLoRA fine-tuned a 65-billion-parameter model on one 48 GB GPU. It used 4-bit base weights and low-rank adapters. Resource efficiency expands access, but it does not repair poor training data or weak evaluation.
QLoRA alcanzó 65B de parámetros en unaGPUTamaños de modelos ajustados en el trabajo de QLoRA; el modelo de 65B usó una GPU de 48 GB con pesosbase de 4 bits.Made for Oscar Ruiz, by HanademiFuentes: Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantizedLLMs. Advances in Neural Information Processing Systems, 36.UnidadMiles de millones de parámetros7B713B1333B3365B65
QLoRA ajustó un modelo de 65.000 millones de parámetros en una GPU de 48 GB. Utilizó pesos base de 4 bits y adaptadores de bajo rango. La eficiencia amplía el acceso, pero no corrige datos deficientes ni evaluaciones débiles.
LoRA made model adaptation dramaticallysmallerIndexed comparison from the paper's GPT-3 example; parameters and memory use separate panels.Made for Oscar Ruiz, by HanademiSources: Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference onLearning Representations.UnitIndexed resource requirementTrainable parameter index10,0001Full tuningLoRAGPU memory index31Full tuningLoRA
LoRA reported about 10,000 times fewer trainable parameters in a GPT-3 example. It also used roughly 3 times less GPU memory than full fine-tuning. Efficient adaptation changes experimentation costs, but it still inherits the quality of the data and evaluation.
LoRA redujo drásticamente el tamaño de laadaptaciónComparación indexada del ejemplo GPT-3 del artículo; los parámetros y la memoria usan panelesseparados.Made for Oscar Ruiz, by HanademiFuentes: Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference onLearning Representations.UnidadRequisito de recursos indexadoÍndice de parámetros entrenables10.0001Ajuste completoLoRAÍndice de memoria GPU31Ajuste completoLoRA
LoRA informó cerca de 10.000 veces menos parámetros entrenables en un ejemplo GPT-3. También utilizó aproximadamente 3 veces menos memoria GPU que el ajuste completo. La adaptación eficiente cambia el coste de experimentar, pero todavía hereda la calidad de los datos y la evaluación.
An open-book system still fails when itretrieves or places evidence badly.The retained data identifies the evaluated tasks but does not contain comparabletask-level effect sizes.Sources: Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33.; Liu, N.F., et al. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157-173.Equal values identify evaluated datasets, not equal performance gains.
The original RAG work reported gains over parametric-only baselines across several knowledge-intensive tasks. Lost in the Middle later found that long-context models often performed worse when relevant evidence appeared in the middle. Retrieval quality and evidence placement are separate engineering problems.
Un sistema con libro abierto todavía falla sirecupera o coloca mal la evidencia.Los datos conservados identifican las tareas evaluadas, pero no contienen tamaños deefecto comparables por tarea.Fuentes: Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33.;Liu, N. F., et al. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157-173.Los valores iguales identifican conjuntos evaluados, no mejoras iguales de rendimiento.
El trabajo original de RAG informó mejoras frente a baselines solo paramétricos en varias tareas intensivas en conocimiento. Lost in the Middle halló después que los modelos de contexto largo rendían a menudo peor cuando la evidencia relevante aparecía en el medio. La calidad de recuperación y la posición de la evidencia son problemas separados.
The original GAIA assistant trailed humansby 77 pointsApproximate GAIA performance for humans and the strongest evaluated assistant configuration.Made for Oscar Ruiz, by HanademiSources: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on LearningRepresentations.; OpenAI. (2024). Introducing SWE-bench Verified. OpenAI.77 percentage points equals 92% minus 15%.UnitPerformanceHumans92%Best evaluated assistant15%6.1x
The original GAIA result reported about 92% human performance and roughly 15% for the strongest evaluated assistant. SWE-bench Verified provides another complete-task approach with 500 human-validated software issues. Agents should be judged by testable outcomes, not supervised demonstrations.
El asistente original de GAIA quedó 77puntos detrás de los humanosRendimiento aproximado en GAIA para humanos y la mejor configuración de asistente evaluada.Made for Oscar Ruiz, by HanademiFuentes: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.;OpenAI. (2024). Introducing SWE-bench Verified. OpenAI.77 puntos porcentuales equivalen a 92% menos 15%.UnidadRendimientoHumanos92 %Mejor asistente evaluado15 %6,1x
El resultado original de GAIA informó cerca de 92% de rendimiento humano y aproximadamente 15% para el mejor asistente evaluado. SWE-bench Verified ofrece otro enfoque de tareas completas con 500 incidencias de software validadas por humanos. Los agentes deben evaluarse por resultados comprobables, no por demostraciones supervisadas.
Treat retrieved text like an untrustedattachment, not a management order.The displayed categories are examples from the full list and are not ordered by severity.Sources: OWASP Foundation. (2025). OWASP Top 10 for Large Language Model Applications 2025. OWASP.Equal values identify risk categories and do not rank severity.
OWASP's 2025 list contains 10 LLM application risk categories. Named examples include prompt injection, sensitive-information disclosure, data poisoning and excessive agency. Permissions, logging, validation and rollback belong in the design from the start.
Trata el texto recuperado como un adjuntono confiable, no como una orden.Las categorías mostradas son ejemplos de la lista completa y no están ordenadas porgravedad.Fuentes: OWASP Foundation. (2025). OWASP Top 10 for Large Language Model Applications 2025. OWASP.Los valores iguales identifican categorías de riesgo y no ordenan su gravedad.
La lista de OWASP de 2025 contiene 10 categorías de riesgos para aplicaciones LLM. Entre ellas aparecen inyección de prompts, divulgación de información sensible, envenenamiento de datos y agencia excesiva. Los permisos, registros, validación y reversión deben estar en el diseño desde el inicio.
React reached near 40% usage, but early AIproducts can use simpler interfaces.A usable control surface matters more than a polished frontend during early validation.Sources: Stack Overflow. (2024). 2024 Developer Survey: Web frameworks and technologies. Stack Overflow.
Stack Overflow's 2024 survey placed React near 40% usage and ahead of several listed alternatives. That makes basic React knowledge useful for exposing AI behavior to users. Early projects should still choose the simplest interface that allows real testing.
React alcanzó cerca de 40% de uso, perolos primeros productos pueden usarinterfaces más simples.Una interfaz utilizable importa más que un frontend pulido durante la validación inicial.Fuentes: Stack Overflow. (2024). 2024 Developer Survey: Web frameworks and technologies. Stack Overflow.
La encuesta de Stack Overflow de 2024 situó React cerca de 40% de uso y por delante de varias alternativas listadas. Eso vuelve útil un conocimiento básico de React para exponer el comportamiento de la IA a usuarios. Los primeros proyectos deben elegir la interfaz más simple que permita pruebas reales.
Whisper evaluated 99 languages and foundlarge language-level error differences.Language count measures coverage, not equal accuracy. The retained claim does notcontain per-language error values.Sources: Radford, A., et al. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of the 40th International Conference onMachine Learning.
Whisper evaluated speech recognition across 99 languages. The paper reported large language-level error differences associated with available training data. Multimodal products therefore need evaluation across the actual languages and input conditions of their users.
Whisper evaluó 99 idiomas y encontrógrandes diferencias de error entre lenguas.El número de idiomas mide cobertura, no precisión igual. La afirmación conservada nocontiene valores de error por idioma.Fuentes: Radford, A., et al. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of the 40th International Conference onMachine Learning.
Whisper evaluó reconocimiento de voz en 99 idiomas. El artículo informó grandes diferencias de error entre lenguas asociadas con los datos de entrenamiento disponibles. Por eso, los productos multimodales necesitan evaluación en los idiomas y condiciones reales de sus usuarios.
Only 1% of surveyed executives called theirgenerative-AI deployments mature.The differentiator is operational judgment: choosing, measuring, securing and maintainingthe smallest system that solves the problem.Sources: McKinsey & Company. (2025). The state of AI: How organizations are rewiring to capture value. McKinsey & Company.
Only 1% of surveyed executives described their generative-AI deployments as mature. That gap makes deployment, evaluation, monitoring, security and maintenance the roadmap's destination. A portfolio should prove that a system remains useful after the first successful prompt.
Solo 1% de los ejecutivos calificósus despliegues de IA generativacomo maduros.El diferenciador es el criterio operativo: elegir, medir, asegurar y mantener el sistemamás pequeño que resuelva el problema.Fuentes: McKinsey & Company. (2025). The state of AI: How organizations are rewiring to capture value. McKinsey & Company.
Solo 1% de los ejecutivos encuestados describió sus despliegues de IA generativa como maduros. Esa brecha convierte despliegue, evaluación, monitoreo, seguridad y mantenimiento en el destino del roadmap. Un portafolio debe demostrar que el sistema sigue siendo útil después del primer prompt exitoso.
In summaryPercent of performance variance attributed to deliberate practice across five domainsMade for Oscar Ruiz, by HanademiSources: Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions. Psychological Science.; Deliberate Practice andPerformance in Music, Games ....; prepnplaced.com.; Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations.; LORA:.; Mialon, G., et al.(2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.; GAIA: a benchmark for general AI assistants | Research.; McKinsey & Company. (2025). The state of AI:…26%Games / Juegos21%Music / Música18%Sports / Deportes4%Education / Educación1%Professions / Profesiones
The roadmap is not a race through every tool. Build foundations, adapt models efficiently, test complete tasks and operate the result safely. Practice alone explained 1% in professions, while the market showed 908 of 15,429 postings at entry or junior level and only 1% of surveyed deployments described as mature. LoRA reported about 10,000 times fewer trainable parameters and roughly three times less GPU memory, while GAIA reported 92% human performance versus roughly 15% for its strongest evaluated assistant configuration.
En resumenPorcentaje de la varianza del rendimiento atribuido a la práctica deliberada en cinco ámbitosMade for Oscar Ruiz, by HanademiFuentes: Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions. Psychological Science.; Deliberate Practice andPerformance in Music, Games ....; prepnplaced.com.; Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations.; LORA:.; Mialon, G., et al.(2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.; GAIA: a benchmark for general AI assistants | Research.; McKinsey & Company. (2025). The state of AI:…26 %Games / Juegos21 %Music / Música18 %Sports / Deportes4 %Educación1 %Professions / Profesiones
El roadmap no es una carrera por todas las herramientas. Construye fundamentos, adapta modelos con eficiencia, prueba tareas completas y opera el resultado con seguridad. La práctica por sí sola explicó 1% en profesiones, mientras el mercado mostró 908 de 15.429 vacantes de entrada o junior y solo 1% de despliegues encuestados descritos como maduros. LoRA informó cerca de 10.000 veces menos parámetros entrenables y aproximadamente tres veces menos memoria GPU, mientras GAIA informó 92% de rendimiento humano frente a aproximadamente 15% para su mejor configuración evaluada.
In summaryMade for Oscar Ruiz, by HanademiSources: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.; Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA:Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36.; Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention withIO-awareness. Advances in Neural Information Processing Systems, 35.; Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations…The original GAIA paper reported about 92% human performance and roughly 15% for the strongest evaluated assistant configuration.QLoRA fine-tuned a 65-billion-parameter model on one 48 GB GPU using 4-bit base weights and low-rank adapters.The FlashAttention paper reported up to threefold faster training and reduced attention-memory growth from quadratic to linear in sequence length.The LoRA paper reported about 10,000 times fewer trainable parameters and roughly three times lower GPU memorythan full fine-tuning in a GPT-3 example.A 2014 meta-analysis attributed 26% of performance variance in games, 21% in music, 18% in sports, 4% in education and1% in professions to deliberate practice.Only 908 of 15,429 Indian tech postings were entry- or junior-level, while 50.1% requested senior engineers in May–July 2026.
The value of the research is not only what each source knew, but what became visible when their evidence was combined.
En resumenMade for Oscar Ruiz, by HanademiFuentes: Mialon, G., et al. (2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations.; Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA:Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36.; Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention withIO-awareness. Advances in Neural Information Processing Systems, 35.; Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations…El artículo original de GAIA informó cerca de 92% de rendimiento humano y aproximadamente 15% para la mejorconfiguración de asistente evaluada.QLoRA ajustó un modelo de 65.000 millones de parámetros en una GPU de 48 GB usando pesos base de 4 bits y adaptadores de bajo rango.El artículo de FlashAttention informó entrenamiento hasta tres veces más rápido y redujo el crecimiento de memoria decuadrático a lineal con la longitud.El artículo de LoRA informó cerca de 10.000 veces menos parámetros entrenables y aproximadamente tres veces menosmemoria GPU que el ajuste completo en un ejemplo GPT-3.Un metaanálisis de 2014 atribuyó a la práctica deliberada 26% de la varianza en juegos, 21% en música, 18% en deportes,4% en educación y 1% en profesiones.Solo 908 de 15.429 vacantes tecnológicas en India eran de entrada o junior, mientras 50,1% buscaba ingenieros seniorentre mayo y julio de 2026.
El valor de la investigación no está solo en cada fuente, sino en lo que apareció al combinar sus evidencias.

The research behind this deck

Hours matter, but time alone cannot certify professional competence. These names describe different layers, not interchangeable skills.

Key findings

The argument

This research is published in English and Spanish. Ver en español

La investigación detrás de esta presentación

Las horas importan, pero el tiempo por sí solo no certifica competencia profesional. Estos nombres describen capas distintas, no habilidades intercambiables.

Hallazgos clave

El argumento

Esta investigación se publica en inglés y español. Read in English