Suche
  • National Library of Norway: Preserving Cultural Heritage for the Digital Future

    National Library of Norway: Preserving Cultural Heritage for the Digital Future

Libraries preserve the past while shaping the future. When cultural memories accumulated over time meet the innovative technologies of the AI era, the wisdom of history finds new ways to inspire the future.

The National Library of Norway has a legal mandate to collect and preserve every publication produced in Norway. Rare books from the 16th century, handwritten newspapers from the 18th century, radio broadcasts and television programs from the 20th century, and digital publications and audiovisual works since 2000 are all preserved here. This collection is an important part of Norway's cultural heritage and a shared legacy for the world. It is an invaluable source of knowledge for generations of scholars, embodying the library's mission to preserve the past for the future.

National Library of Norway

Over 20 Years of Digitalization

The National Library of Norway has been undertaking a large-scale digitization program for over two decades, one of the longest-running digitization efforts in Europe. However, this is a complex process that involves far more than simply scanning and archiving files.

Each item goes through a sophisticated processing pipeline, beginning with the capture of physical materials. This entails scanning book pages, photographing objects, and recording audio and video. The content then undergoes extensive post-processing, including image correction, format normalization, and quality control. In the next stage, optical character recognition (OCR) is used to extract text from millions of pages, including historical typefaces, old Norwegian dialects, and even manuscripts dating back to the 19th century. In the final stage, metadata tagging, language detection, and structural analysis are performed.

A single 25-page newspaper issue can generate around 150 independent files, including master scans, derivative copies, OCR output, and metadata. All of these files are packaged into a single archive object. The scale becomes truly staggering when this is multiplied for millions of objects and for over 20 years.

The National Library of Norway has amassed the world's largest, oldest, and most comprehensive Norwegian-language corpus. It extends beyond what search engines such as Google can index.

From Archived Data to AI-Ready Assets

Since the launch of the program, about 20 petabytes (PB) of unique digital objects have been collected. To ensure no data is ever lost, the collection is retained under a 3+2+1 data protection policy: three copies across two storage technologies (tapes and disks), with one copy stored offsite. Together, they occupy approximately 60 PB of storage.

In recent years, as artificial intelligence (AI) sweeps over the world, the Norwegian government has identified a critical challenge. Norwegian has only around five million speakers, but it comprises two official written standards and dozens of dialects. The language was significantly underserved by leading AI models, which are trained predominantly on English. Without developing its own large language models (LLMs), Norway would have to rely on AI systems built elsewhere—systems that may never fully understand Norway's language, culture, and social context. A country that cannot interact with AI in its own language on equal terms will inevitably be at a disadvantage in the AI era.

The Norwegian Ministry of Culture and Equality responded by entrusting the library with developing the country's own LLMs, backed by an annual allocation of NOK70 million. The library also benefits from a legal agreement built over decades with Norwegian newspapers and publishers that allows it to train models on copyrighted content. This is an institutional advantage beyond the reach of private companies.

Striking the Right Balance with All-Flash Storage

Unlocking new technology applications from historical data presents a seemingly contradictory challenge for data infrastructure. The National Library of Norway must simultaneously address two fundamentally different storage needs.

The first is long-term archiving, where durability and cost efficiency are paramount. Nearly 60 PB of mass data is stored on tapes and hard disk drives (HDDs). Performance is not the primary concern; instead, the data must remain accessible and readable after a thousand years.

The second need is the opposite. AI model training is highly dependent on adequate data preparation. Once preprocessing begins, massive volumes of historical data must be retrieved as quickly as possible, as any latency bottleneck means wasted compute resources. However, migrating petabytes of data from the long-term archiving system to a new storage environment is a significant challenge.

This is where Huawei's OceanStor Dorado All-Flash Storage comes into play. Deployed within the National Library of Norway's on-premises AI environment, the storage system provides a total capacity of more than 2 PB, which is dedicated to extracting data from the archive and performing data cleansing, deduplication, formatting, and validation. With ultra-low latency and high throughput, Huawei OceanStor Dorado All-Flash Storage brings valuable historical and cultural data online. After being processed by CPU clusters, the prepared datasets are efficiently delivered to Sigma2, Norway's national e-infrastructure, for final model training.

The National Library of Norway has long recognized the importance of high-performance storage. From its early SAS architecture to today's NVMe all-flash architecture, throughput and latency have consistently been the defining metrics for the library's infrastructure. Without high-performance innovative storage solutions as a bridge, the gap between archiving and computing would be difficult to overcome.

The Journey Continues

The National Library of Norway has already made great strides in its AI initiatives. Its speech recognition and text generation tools built on an open-source architecture are now being adopted in different sectors in Norway.

Looking ahead, opportunities go hand in hand with new challenges. For example, how can Norway establish a Norwegian (with two written forms and various dialects) model quality evaluation system? How should governance frameworks for national AI tools be designed to ensure that LLMs serve the public interest in a fair, resilient, and trustworthy manner? These questions will be addressed through continual exploration and practical application.

Preserving the Past to Build the Future

The National Library of Norway's journey toward digital intelligence raises a thought-provoking question: What is the most valuable infrastructure in the AI era?

The answer is not computing power. It is data that has been carefully collected, organized, and preserved long before AI existed, enabled by a steadfast commitment to preserving cultural heritage. These carefully protected records are the foundation on which innovative technologies can reflect a nation's language, culture, and history.

"We started digitizing to preserve Norway's past for the future. We ended up building the infrastructure for Norway's AI future."

Marius Husnes

Head of IT Platform at the National Library of Norway

TOP