MillenniumDB: the powerful multimodal engine created at the Millennium Institute Foundational Research on Data (IMFD)

December, 2023.- In a world where an enormous amount of data is produced every second: how can we extract information that is useful to us? This is the question that researchers at the Millennium Institute Foundational Research on Data (IMFD) have been working to answer for over a decade.

“Currently, Millennium DB is a data management system that allows knowledge graphs containing a very large volume of data to be handled efficiently, and it effectively uses the most modern techniques that exist and are known from research,” explains Domagoj Vrgoč, academic at the Institute of Mathematical and Computational Engineering of the Universidad Católica de Chile, researcher at the Millennium Institute Foundational Research on Data (IMFD) and one of the lead authors of the research behind this work.

Domagoj Vrgoč

Millennium DB is a modular, open-source data management engine that allows a variety of knowledge graphs storing a very large volume of data to be handled efficiently. Millennium DB uses techniques that are at the forefront of scientific research in this field at the moment: it is based on a combination of proven data management techniques, state-of-the-art algorithms for worst-case-optimal joins, as well as specialized algorithms for evaluating path queries. It allows different data formats to be combined and graphs to be created, from which useful information can be obtained in different forms.

This is software that was developed from scratch in Chile, and it has the capacity to compete with, and even outperform, other systems that have been in development for years. In general, all these kinds of tools are created in the global north: with Millennium DB we can say that in Chile we carry out and develop quality research that can compete with developments from other countries,” explains Carlos Rojas, researcher at the Millennium Institute Foundational Research on Data (IMFD) and director of the project.

Millennium DB has already been tested with data from different fields; the first tests were carried out with Wikidata and “one example where we used this model was for the analysis carried out on the country’s social and political situation during the period when the Constitutional Convention was operating,” explains Juan Reutter, deputy director of the IMFD. “Beyond the outcome of the Convention, during that period we had a large amount of information: comments on social media, the complete broadcasts of the debates held in the former congress, surveys that we also conducted with different audiences and with temporal components, and the very text that was being created: all of that we were able to systematize in this model and obtain very enriching analyses that allow us to combine variables of different types.”

Juan Reutter

For Jazmine Maldonado, Director of Innovation and Technology Transfer at the Millennium Institute Foundational Research on Data (IMFD), the model has great potential to be developed as a product that can boost different sectors of the market. “For example, let’s think about retail: you have a series of information obtained from customers, such as purchasing preferences, the searches they carry out, the periods when something specific is searched for, and you also have product prices, features, perhaps images, codes. Companies have all this information available, but it is difficult to find a way to systematize it that allows them to obtain data useful for the core of their business.”