Organizada por semana. Las lecturas centrales están marcadas con ●. Los enlaces llevan a versiones de acceso abierto cuando existen; las ediciones en español se indican para que el estudiante pueda buscarlas en biblioteca.
Fuentes de acceso abierto que cubren buena parte del curso
- Python Software Foundation. Documentación de Python en español: tutorial, biblioteca estándar y guías. https://docs.python.org/es/3/
- pandas development team. Documentación de pandas. https://pandas.pydata.org/docs/
- The Programming Historian en español. Lecciones revisadas por pares sobre métodos digitales para humanidades. https://programminghistorian.org/es/
- Secretaría del Senado de Colombia. Textos normativos consolidados. http://www.secretariasenado.gov.co/
- Gobierno de Colombia. Portal de datos abiertos. https://www.datos.gov.co/
- Biblioteca Nacional de Colombia. Colecciones digitales y hemeroteca. https://bibliotecanacional.gov.co/
- ACL Anthology. Literatura de lingüística computacional citada en el curso. https://aclanthology.org/
- arXiv. Literatura técnica sobre anotación con modelos de lenguaje. https://arxiv.org/
Semana 1. Qué es un corpus y qué lo hace válido
- ● Valles, M. S. (1997). Técnicas cualitativas de investigación social. Reflexión metodológica y práctica profesional. Madrid: Síntesis. Cap. 4.
- ● Nguyen, D., Liakata, M., DeDeo, S., Eisenstein, J., Mimno, D., Tromble, R. y Winters, J. (2020). "How We Do Things With Words: Analyzing Text as Social and Cultural Data". Frontiers in Artificial Intelligence 3, 62. https://doi.org/10.3389/frai.2020.00062
- Scott, J. (1990). A Matter of Record: Documentary Sources in Social Research. Cambridge: Polity Press.
- Biber, D. (1993). "Representativeness in Corpus Design". Literary and Linguistic Computing 8(4), 243-257. https://doi.org/10.1093/llc/8.4.243
- Glaser, B. G. y Strauss, A. L. (1967). The Discovery of Grounded Theory: Strategies for Qualitative Research. Chicago: Aldine. Cap. 3, sobre muestreo teórico.
- Flick, U. (2015). El diseño de la investigación cualitativa. Madrid: Morata.
- Flick, U. (2007). Introducción a la investigación cualitativa. 2.ª ed. Madrid: Morata.
- Congreso de Colombia. Ley 23 de 1982, sobre derechos de autor. http://www.secretariasenado.gov.co/senado/basedoc/ley_0023_1982.html
- Congreso de Colombia. Ley 1581 de 2012, protección de datos personales. http://www.secretariasenado.gov.co/senado/basedoc/ley_1581_2012.html
- Congreso de Colombia. Ley 1712 de 2014, transparencia y acceso a la información pública. http://www.secretariasenado.gov.co/senado/basedoc/ley_1712_2014.html
Semana 2. Python mínimo para humanidades
- ● Python Software Foundation. El tutorial de Python. Caps. 3 y 7. https://docs.python.org/es/3/tutorial/
- ● Rockwell, G. y Sinclair, S. (2016). Hermeneutica: Computer-Assisted Interpretation in the Humanities. Cambridge, MA: MIT Press. Cap. 1.
- Karsdorp, F., Kestemont, M. y Riddell, A. (2021). Humanities Data Analysis: Case Studies with Python. Princeton: Princeton University Press. Acceso abierto: https://www.humanitiesdataanalysis.org/
- Project Jupyter. Documentación de JupyterLab. https://jupyterlab.readthedocs.io/
- Python Software Foundation. Módulos
pathlibycsv. https://docs.python.org/es/3/library/pathlib.html y https://docs.python.org/es/3/library/csv.html
Semana 3. Recolección legal y reproducible
- ● Congreso de Colombia. Ley 1712 de 2014. Títulos I y II. http://www.secretariasenado.gov.co/senado/basedoc/ley_1712_2014.html
- ● Congreso de Colombia. Ley 1581 de 2012. Títulos I a III. http://www.secretariasenado.gov.co/senado/basedoc/ley_1581_2012.html
- ● Koster, M., Illyes, G., Zeller, H. y Sassman, L. (2022). Robots Exclusion Protocol. RFC 9309. Internet Engineering Task Force. https://www.rfc-editor.org/rfc/rfc9309
- Reitz, K. y colaboradores. Requests: HTTP for Humans. https://requests.readthedocs.io/
- Richardson, L. Beautiful Soup Documentation. https://www.crummy.com/software/BeautifulSoup/bs4/doc/
- Python Software Foundation. Módulos
hashlib,timeyjson. https://docs.python.org/es/3/library/hashlib.html
Semana 4. El catálogo maestro
- ● Gibbs, G. (2012). El análisis de datos cualitativos en investigación cualitativa. Madrid: Morata. Cap. 3.
- ● Wilkinson, M. D. et al. (2016). "The FAIR Guiding Principles for scientific data management and stewardship". Scientific Data 3, 160018. https://doi.org/10.1038/sdata.2016.18
- pandas development team. 10 minutes to pandas. https://pandas.pydata.org/docs/user_guide/10min.html
- pandas development team. Merge, join, concatenate and compare. https://pandas.pydata.org/docs/user_guide/merging.html
- Gebru, T., Morgenstern, J., Vecchione, B., Wortman Vaughan, J., Wallach, H., Daumé III, H. y Crawford, K. (2021). "Datasheets for Datasets". Communications of the ACM 64(12), 86-92. https://arxiv.org/abs/1803.09010
- Gitelman, L. (ed.) (2013). "Raw Data" Is an Oxymoron. Cambridge, MA: MIT Press.
Semana 5. Limpieza y normalización
- ● Rawson, K. y Muñoz, T. (2019). "Against Cleaning". En M. K. Gold y L. F. Klein (eds.), Debates in the Digital Humanities 2019. Minneapolis: University of Minnesota Press. https://dhdebates.gc.cuny.edu/
- ● Cordell, R. (2017). "'Q i-jtb the Raven': Taking Dirty OCR Seriously". Book History 20, 188-225. https://doi.org/10.1353/bh.2017.0006
- Python Software Foundation. Unicode HOWTO. https://docs.python.org/es/3/howto/unicode.html
- Python Software Foundation. Módulos
unicodedatayre. https://docs.python.org/es/3/library/unicodedata.html y https://docs.python.org/es/3/library/re.html - Drucker, J. (2011). "Humanities Approaches to Graphical Display". Digital Humanities Quarterly 5(1). http://www.digitalhumanities.org/dhq/vol/5/1/000091/000091.html
Semana 6. Primera lectura computacional
- ● Sinclair, J. (1991). Corpus, Concordance, Collocation. Oxford: Oxford University Press. Caps. 1 y 8.
- ● Moretti, F. (2015). Lectura distante. Buenos Aires: Fondo de Cultura Económica. Cap. 1. Original: Distant Reading, Londres: Verso, 2013.
- Moretti, F. (2000). "Conjectures on World Literature". New Left Review 1, 54-68.
- Rockwell, G. y Sinclair, S. (2016). Hermeneutica. Caps. 2 y 3. Voyant Tools: https://voyant-tools.org/
- Firth, J. R. (1957). "A Synopsis of Linguistic Theory, 1930-1955". En Studies in Linguistic Analysis. Oxford: Blackwell, 1-32.
- Church, K. W. y Hanks, P. (1990). "Word Association Norms, Mutual Information, and Lexicography". Computational Linguistics 16(1), 22-29. https://aclanthology.org/J90-1003/
- Manning, C. D. y Schütze, H. (1999). Foundations of Statistical Natural Language Processing. Cambridge, MA: MIT Press. Cap. 5.
- Zipf, G. K. (1949). Human Behavior and the Principle of Least Effort. Cambridge, MA: Addison-Wesley.
- Jockers, M. L. (2013). Macroanalysis: Digital Methods and Literary History. Urbana: University of Illinois Press.
- Underwood, T. (2019). Distant Horizons: Digital Evidence and Literary Change. Chicago: University of Chicago Press.
- Ramsay, S. (2011). Reading Machines: Toward an Algorithmic Criticism. Urbana: University of Illinois Press.
- Python Software Foundation.
collections.Counter. https://docs.python.org/es/3/library/collections.html
Semana 7. Lectura asistida y registro de frontera
- ● Grimmer, J. y Stewart, B. M. (2013). "Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts". Political Analysis 21(3), 267-297. https://doi.org/10.1093/pan/mps028
- ● Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z. y Yang, D. (2024). "Can Large Language Models Transform Computational Social Science?". Computational Linguistics 50(1), 237-291. https://doi.org/10.1162/coli_a_00502
- Grimmer, J., Roberts, M. E. y Stewart, B. M. (2022). Text as Data: A New Framework for Machine Learning and the Social Sciences. Princeton: Princeton University Press.
- Gilardi, F., Alizadeh, M. y Kubli, M. (2023). "ChatGPT outperforms crowd workers for text-annotation tasks". Proceedings of the National Academy of Sciences 120(30). https://doi.org/10.1073/pnas.2305016120
- Pangakis, N., Wolken, S. y Fasching, N. (2023). "Automated Annotation with Generative AI Requires Validation". https://arxiv.org/abs/2306.00176
- Krippendorff, K. (1990). Metodología de análisis de contenido. Teoría y práctica. Barcelona: Paidós.
- Bardin, L. (1986). El análisis de contenido. Madrid: Akal.
- Zainea, C. I. (2026). "Los modelos de lenguaje no alucinan: cometen infortunios". https://izainea.github.io/blog/alucinacion-o-infortunio/
- Zainea, C. I. (2026). "La prótesis hermenéutica". https://izainea.github.io/blog/protesis-hermeneutica/
Semana 8. El corpus como aserción
- ● Bender, E. M. y Friedman, B. (2018). "Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science". Transactions of the Association for Computational Linguistics 6, 587-604. https://aclanthology.org/Q18-1041/
- ● Zainea, C. I. (2026). Agencia sin imputación. Hermenéutica, pragmática y la validación epistémica del discurso artificial. https://izainea.github.io/escritos/agencia-sin-imputacion.pdf
- Gebru, T. et al. (2021). "Datasheets for Datasets". Communications of the ACM 64(12), 86-92. https://arxiv.org/abs/1803.09010
- Wilkinson, M. D. et al. (2016). "The FAIR Guiding Principles". Scientific Data 3, 160018. https://doi.org/10.1038/sdata.2016.18
- Nguyen, D. et al. (2020). "How We Do Things With Words". Frontiers in Artificial Intelligence 3, 62. https://doi.org/10.3389/frai.2020.00062
- Zainea, C. I. (2026). "Agencia sin imputación: por qué «la máquina no comprende» no cierra la discusión". https://izainea.github.io/blog/agencia-sin-imputacion/
- Python Packaging Authority. pip freeze. https://pip.pypa.io/en/stable/cli/pip_freeze/