Image and Research Conference
Until recently, computational research in the social sciences and humanities worked in a single modality, analysing either text or images in isolation. Recent advances in multimodal machine learning have changed this. Trained on image-text combinations, multimodal AI can connect images to texts and texts to images. This talk shows what that shift makes possible through one extended case study. Martí Massafont Costals (1918–2012) worked the streets of Girona as a commercial photographer for five decades, leaving 35,000 negatives, of which 9,021 have been digitised. Scholars have largely approached the collection through close reading of individual images. In this presentation, I show how a clustering algorithm can group all 9,021 photographs into 78 fine-grained clusters. These groups of images show the kinds of people, scenes, and places Massafont photographed, revealing a commercial grammar of purchasable moments: a shared understanding between Massafont and his customers about which occasions were worth commemorating and how they should be framed. The photographs the algorithm could not cluster, classified as 'noise', mark the edges of that grammar. I argue that this kind of distant viewing of a visual archive can help historians find patterns of presence and patterns of absence alike.
Thomas Smits is Assistant Professor of Digital History & AI at the University of Amsterdam, where he co-directs the Amsterdam Computational History Lab. He studies how images have shaped public understanding of the world since the nineteenth century. He is currently using AI to read colonial aerial photographs of Indonesia, and developing multimodal AI methods that cluster large photographic collections to surface visual patterns no catalogue records.