Skip to content. | Skip to navigation

Personal tools
You are here: Home Publications Constance: An Intelligent Data Lake System


Prof. Dr. S. Decker
RWTH Aachen
Informatik 5
Ahornstr. 55
D-52056 Aachen
Tel +49/241/8021501
Fax +49/241/8022321

How to find us

Annual Reports





Constance: An Intelligent Data Lake System

Year 2016
Abstract URL view
ISBN-13 978-1-4503-3531-7

As the challenge of our time, Big Data still has many research hassles, especially the variety of data. The high diversity of data sources often results in information silos, a collection of non-integrated data management systems with heterogeneous schemas, query languages, and APIs. Data Lake systems have been proposed as a solution to this problem, by providing a schema-less repository for raw data with a common access interface. However, just dumping all data into a data lake without any metadata management, would only lead to a `data swamp'. To avoid this, we propose Constance, a Data Lake system with sophisticated metadata management over raw data extracted from heterogeneous data sources. Constance discovers, extracts, and summarizes the structural metadata from the data sources, and annotates data and metadata with semantic information to avoid ambiguities. With embedded query rewriting engines supporting structured data and semi-structured data, Constance provides users a unified interface for query processing and data exploration. During the demo, we will walk through each functional component of Constance. Constance will be applied to two real-life use cases in order to show attendees the importance and usefulness of our generic and extensible data lake system.


In Proceedings of the 2016 International Conference on Management of Data (SIGMOD), pp. 2097-2100. ACM.

Presented at

ACM SIGMOD, 2016 , San Francisco , US.

Published in

SIGMOD '16 Proceedings of the 2016 International Conference on Management of Data , p. 2097-2100 ; ACM , New York , US .

Related projects

Document Actions