Towards the understanding of hierarchical structures in newspaper images

Two AI approaches to understanding hierarchy in newspaper images: a modular pipeline and the Tiramisu transformer. Discover its advantages!

domingo, 19 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Two approaches: modular pipeline and Tiramisu transformer

The digitization of historical newspapers represents a fascinating technical challenge, as these documents present complex hierarchical structures: sections, articles, text blocks, images and captions are organized in dense and heterogeneous layouts. Automatically understanding this hierarchy not only allows for the search and retrieval of information, but opens the door to artificial intelligence applications that can extract valuable knowledge for historians, journalists, and analysts. In this article, we explore the most advanced approaches to understanding hierarchical structures in newspaper images, from modular pipelines to transformer-based end-to-end architectures, and discuss how these technologies can be integrated into enterprise solutions.

The main challenge is that a newspaper is not just a collection of paragraphs: it contains titles, subheadings, columns, advertisements, images with footers, and elements that are nested within sections. Traditional OCR (optical character recognition) systems process text in a linear way, but lose the semantic structure. To overcome this, two major trends have been developed: the first is committed to a bottom-up modular approach, combining layout detection models, reading order prediction and article segmentation. The second proposes novel architectures that model hierarchy in an iterative way, using parallelized attention mechanisms. Both ways are driving the creation of tailor-made applications for the digitization of heritage and the automation of document processes.

In the modular approach, detectors such as YOLO are employed to locate regions (titles, images, columns), then a reading order model (such as LayoutReader) assigns a logical sequence, and finally a custom algorithm groups those blocks into articles. The advantage is flexibility: each component can be replaced or improved independently, making it easy to integrate into custom software systems that require customization according to the type of newspaper or language. However, this approach can accumulate errors if a module fails, and article segmentation remains a critical point when texts are cut between columns or pages.

The second current, represented by architectures such as Tiramisu (Tiered Transformers for Hierarchical Structure Understanding), proposes a unified model that processes the entire image and simultaneously generates section separation, block location, semantic categorization, and reading order. It uses transformers at levels that iteratively refine hierarchical understanding. These types of systems are ideal for companies looking for AI with high levels of accuracy and scalability, as they can be trained with synthetic data and then fine-tuned with real collections. In addition, the parallelization of attention mechanisms allows large volumes of images to be processed in cloud environments, such as AWS and Azure cloud services, reducing computing times.

A fundamental aspect in this field is to have adequate datasets for evaluation. The release of collections specifically tagged for hierarchical retrieval, such as the Finlam La Liberté dataset, allows methods to be compared and research to be advanced. For companies, having this data is a first step in developing AI agents that automate the indexing of documentary collections, making it easier for historians or citizens to consult them. In fact, combining these models with business intelligence tools such as power bi can transform data extracted from newspapers into interactive dashboards that show historical trends, mentions of characters or evolution of topics.

In addition, cybersecurity plays an important role in the management of these systems: digitized newspaper repositories often contain sensitive or patrimonial information protected by copyright, so it is necessary to implement access and auditing controls. At Q2BSTUDIO, as a software and technology development company, we offer bespoke applications that integrate AI models, cloud processing pipelines, and security layers, adapting to the needs of museums, libraries, and media companies. Our team has worked on digitization projects where AWS and Azure cloud services are combined with computer vision models, managing to extract hierarchical structures from historical documents with high fidelity.

From a practical perspective, implementing a system of understanding hierarchical structures in newspapers requires first defining the scope: is it only the segmentation of articles or also the classification by sections (politics, sports, culture)? Is it necessary to preserve the original reading order? Will you work with low-quality images or modern scans? These questions determine the choice of technical approach. For large-scale projects, we recommend a modular pipeline with the possibility of periodic retraining, while for specific tasks with controlled volumes, an end-to-end architecture can deliver better results. In any case, integration with business intelligence platforms allows end users to not only search, but visualize patterns.

The future of this technology lies in the incorporation of multimodal models that combine text, image and metadata, and in the use of AI agents capable of reasoning about the structure of the document to answer complex questions ("what articles on economics appeared on the cover of 1920?"). At Q2BSTUDIO we are developing AI solutions for companies that integrate these advances, offering consulting, implementation and maintenance services for document knowledge extraction systems. Our expertise in services, business intelligence and bespoke software ensures that any project is tailored to the client's specific requirements.

In conclusion, understanding hierarchical structures in newspaper images is a vibrant field that combines computer vision, natural language processing, and deep learning. Current approaches, both modular and unified, offer powerful tools for digitizing and making documentary heritage accessible. Companies that invest in these solutions not only preserve history, but unlock structured data for analysis and decision-making. With the support of technologists like Q2BSTUDIO, it is possible to build robust, scalable, and secure systems that transform stacks of images into actionable knowledge.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.