Thesis Description & Objectives
Objective: Design and develop AI systems capable of automatically generating coherent natural language or visual summaries from clusters of semantically related images from social media
Context: The automatic summarization of image collections represents a challenging and emerging frontier in multimodal AI. Unlike single-image captioning, cluster-level summarization requires a system to identify recurring themes, abstract away from redundant visual details, and produce a concise, human-readable description that captures the collective content of a group of images. The goal of this project is to design generative models that can generate meaningful summaries (textual and/or visual) at the cluster level. Key challenges are enabling the model to perform cross-image reasoning — identifying what is shared, what varies, and what is most salient across the set — and incorporating existing semiotic models developed by our group.
Activities: The thesis will include one or more of the following activities:
– starting from an analysis of the SOTA [2], design an effective pooling or aggregation strategy that compresses visual information across a cluster into a unified representation suitable for language generation, tailoring the tool to social media listening, market research and cultural studies;
– leverage Large Language Models (e.g., GPT-4, LLaMA) to produce fluent and abstractive summaries grounded in visual content;
– incorporate structured knowledge, also coming from computational semiotics tools such as FRESCO [2], to improve the coherence and factual accuracy of generated summaries;
– Investigate the application of cluster-level summarization to produce AI “personas” based on clusters of similar social media images.
Reading material:
[1] https://arxiv.org/abs/2503.19361
[2] https://arxiv.org/abs/2407.03268

