In traditional IT, one stored data in a way that facilitated the way one planned to use it. For analytics, this meant creating a well-defined and rigidly structured data warehouse based on the types of queries expected. But today, we collect a wide variety of data before we know how we will use them, and value flexibility and agility in their use. A data lake is a collection of data in all of its varieties and forms which can be used flexibly with the broad range of analytic tools now available. In other words, data we collect can be dumped into the data lake with little or no attempt to manipulate them for analysis because today’s analytic tools are more flexible. For example, if your company acquired another company with different IT systems, it used to take time to integrate systems to be able to combine their data for reporting. But with a data lake, you pour in the data from both companies’ systems and can often begin combining them immediately (for analysis, not necessarily for transaction processing).
— War and Peace and IT: Glossary, Mark Schwartz