This is not the first time we have talked about machine learning in ArqueoTimes (Luengo, 2023), but perhaps it is the first time that we focus on such a specific use as its application in epigraphy. As we hinted at in that previous article, and citing Humby and Palmer (2006):
«Data is the new oil. It’s valuable, but if unrefined it cannot really be used. It has to be changed into gas, plastic, chemicals, etc., to create a valuable entity that drives profitable activity; so must data be broken down, analyzed for it to have value» (Data is the new oil. It is valuable, but if not refined, it cannot actually be used. It needs to be converted into gasoline, plastics, chemicals, etc., to create a valuable entity that drives profitable activity; thus, data must be decomposed and analyzed to have value).

In the context of epigraphy, this quote from Humby and Palmer takes on enormous relevance. Epigraphy constantly faces challenges such as the fragmentation or deterioration of inscriptions, and it is here where the Ithaca algorithm developed in Google DeepMind's labs by a team led by Yannis Assael comes into play, providing an advanced tool for restoring and analyzing ancient inscriptions, even when they are damaged or incomplete. However, it is important to note that Ithaca is not the only nor the first example of using machine learning in epigraphy or on ancient texts: Kang et al. (2021) employed neural language models and machine translation techniques to restore and analyze records from the Joseon dynasty. Others like Bamman and Burns (2021) developed Latin BERT, a contextual language model specifically trained for Latin, capable of predicting missing text demonstrating the wide applicability of these technologies in preserving and studying historical texts across various cultural contexts.
But what is Ithaca? Description of the algorithm
Imagine that we have a puzzle where pieces are damaged, worn out, or even completely missing. The work of epigraphers in this context is like that of a puzzle restorer: they try to fit together the remaining pieces and reconstruct the original image based on clues provided by the context. The Ithaca algorithm acts as an add-on that, instead of simply trying to fit visible pieces, can infer the shape and content of missing pieces based on patterns and experiences learned from thousands of similar puzzles (in this case, ancient inscriptions).
The team emphasizes that Ithaca is not meant to replace the scientist but to serve as a powerful support tool. In fact, the study highlights that while an historian alone achieves 25% accuracy and the Ithaca algorithm separately reaches 62%, the combination of both, historian and algorithm, achieves a remarkable 72% accuracy.
But we cannot stop at these data without delving deeper because, what does a 72% accuracy mean in this context, how does this algorithm really work, and why is it revolutionizing epigraphy?
Data and Model of the Ithaca Algorithm
To understand how Ithaca achieves its level of accuracy, it is crucial to understand both the data used in training as well as the architecture of the algorithm. The work, presented to the scientific community on March 9, 2022, in Nature (Assael et al., 2022), introduced Ithaca as a deep neural network designed to help transcribe fragments of text previously illegible and identify the original location and dating of ancient inscriptions, particularly from ancient Greece. To train the model, researchers started with an epigraphic corpus composed of 178,551 inscriptions from the Packard Humanities Institute (PHI), of which 78,608 were selected for training, becoming the largest dataset of epigraphic text managed by machine learning to date. This set of inscriptions is written in ancient Greek and have been found throughout the ancient Mediterranean, with dates ranging from the 7th century BC to the 5th century AD.

The design of the algorithm is based on a transformer architecture (a transformer is a deep learning architecture developed by Google researchers and based on the attention mechanism, according to Vaswani et al., 2017). This system is structured into two main parts: a «body» of transformers and three specific «heads» for each task: text restoration, geographic attribution, and chronological attribution. This approach allows the algorithm to consider context broadly and connections between different parts of the text, even if some parts are fragmented or damaged. To better understand these concepts, we will use the following analogy: think of Ithaca's main body as a person at a party listening to multiple conversations simultaneously. This person can focus on the most relevant conversation while still picking up fragments of nearby conversations. Similarly, Ithaca's transformer focuses on certain parts of the text (words and characters) that are more relevant for understanding the complete context, while still «listening» to the rest of the text that might be relevant for future interpretations. This approach allows the algorithm to consider broad context and connections between different parts of the text, even if some parts are fragmented or damaged.
As epigraphic inputs, Ithaca uses combined character and word representations, addressing thus the loss of parts of the text. Additionally, it introduces a special symbol «[unk]» when there are damaged or unknown words. The outputs from the body are channeled to the three heads, each consisting of a shallow feedforward neural network (shallow FNN) trained specifically for its task. For example, the restoration head predicts missing characters, the geographic attribution head classifies the inscription among 84 regions, and the chronological attribution head estimates the date in intervals of ten years between 800 BC and AD 800. The model also generates interpretative visualizations such as saliency maps and ranked prediction lists that facilitate collaboration between the model and historians.
Regarding the idea of 72% accuracy, in the context of the Ithaca model, this means that in 72% of cases, the highest prediction by the model is correct. As we indicated above, Ithaca is not the first algorithm to perform these tasks. Pythia, another machine learning model developed by a team also led by Assael (2019), has addressed these tasks previously and forms an integral part of the new algorithm. But as expected, Ithaca improves performance in all areas compared to Pythia. For example, for text restoration (Restoration), Ithaca presents a lower character error rate (CER) with 18.3% versus 47.0% of Pythia. Similarly, in the Top-1 option, Ithaca (independently) is far superior, boasting 61.8% against 32.6% for Pythia.

Surely over time we can digitize and enrich epigraphic corpora, increasing the number of training data which will allow us to improve models in the long run. But regarding this and to conclude, perhaps it is appropriate for us to make a fundamental reflection. The accuracy of these models is measured relative to what has been previously labeled. That is, researchers label the training set. Once the model is trained, the algorithm can continue inferring dates on its own. However, if the training data were erroneous, all training would generate an erroneous model that could only offer incorrect answers. In conclusion, and paraphrasing Einstein's ideas about science: «A theory may fall with new evidence, but a wrong datum may become deeply rooted in the fabric of knowledge» (inspired by «My Ideas and Vision of the World», Albert Einstein (2023)).
Bibliography
Assael, Y., Sommerschield, T., Shillingford, B., Bordbar, M., Pavlopoulos, J., Chatzipanagiotou, M., Androutsopoulos, I., Prag, J., & de Freitas, N. (2022). Restoring and attributing ancient texts using deep neural networks. Nature, 603(7900), 280-283. https://doi.org/10.1038/s41586-022-04448-z
Assael, Y., Sommerschield, T., & Prag, J. (2019). Restoring ancient text using deep learning: A case study on Greek epigraphy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 6368-6375). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1668
Bamman, D., & Burns, P. J. (2021). Latin BERT: A contextual language model for classical philology. Manuscript in preparation. https://doi.org/10.48550/arXiv.2009.10053
Bamman, D. (2021). Latin BERT [GitHub Repository]. GitHub. https://github.com/dbamman/latin-bert (accessed on September 5, 2024).
Einstein, A. (2023). My Ideas and Vision of the World (C. Fosch, Illus.; J. M. Álvarez Flórez & A. Goldar, Trans.). Alma.
Humby, C., & Palmer, M. (November 3, 2006). Data is the New Oil. https://ana.blogs.com/maestros/2006/11/data_is_the_new.html (accessed on August 25, 2024).
Kang, K., et al. (2021). Restoring and mining the records of the Joseon dynasty via neural language modeling and machine translation. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL) (pp. 4031–4042). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.320
Luengo Gutiérrez, F. J. (2023). Artificial intelligence (AI) in archaeology. Arqueotimes. /articles/artificial-intelligence-ai-in-the-archaeology/
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (pp. 6000–6010). Curran Associates Inc.
Comments