From Data Lakes to Community Wisdom: Why Information Science Matters More Than Storage

When people hear the term data lake, they often imagine a giant digital warehouse where information is collected and stored. While that description is technically correct, it misses the most important point.

A data lake is not valuable because of what it stores.

A data lake is valuable because of what it helps people learn.

This distinction matters because organizations across education, healthcare, workforce development, economic development, and government are generating more data than ever before. The challenge is no longer collecting information. The challenge is turning that information into knowledge that communities can use to make better decisions.

For initiatives such as MiGreatDataLake, the goal is not to build a bigger database. The goal is to build a learning system.

The DIKW Framework

Information scientists often describe this journey as a progression from:

Data → Information → Knowledge → Wisdom

Each layer answers a different question.

Data: What Happened?

Data consists of raw observations.

Examples include:

  • A student missed school today.
  • A broadband speed test measured 75 Mbps.
  • A workforce training program served 25 participants.
  • A weather sensor recorded four inches of rainfall.

Individually, these observations tell us very little.

They are facts without context.

In a modern cloud architecture, this raw information typically lands in a storage environment such as Amazon S3, where data from many different systems can be stored together in its original form.

At this stage, we have data, but we do not yet have understanding.

Information: What Does It Mean?

Information emerges when we organize data and establish context.

Now we begin asking questions such as:

  • Which school district generated this data?
  • Which county does it belong to?
  • What year was it collected?
  • Are two systems using the same definition for a particular term?

This is where information science begins to create value.

Many data projects fail because organizations assume everyone interprets data in the same way. In reality, different systems often use different terminology, definitions, and assumptions.

One system may report “attendance events.”

Another may report “student absences.”

Are they describing the same thing?

The answer is not a technology question.

It is an information science question.

At this layer, technologies such as AWS Glue help catalog, classify, and organize data. More importantly, information scientists help develop the shared vocabulary needed to ensure that data from multiple sources can be understood consistently.

Without this work, a data lake quickly becomes a data swamp.

Knowledge: Why Is It Happening?

This is where things become truly interesting.

Information tells us what happened.

Knowledge helps us understand why.

Instead of looking at individual data points, we begin identifying patterns, connections, and relationships.

For example, a school district might examine relationships among:

  • Attendance
  • Academic performance
  • Career and Technical Education participation
  • Broadband access
  • Postsecondary enrollment

A traditional reporting system can tell us these variables exist.

Knowledge emerges when we begin understanding how they influence one another.

Historically, this has been the role of analytics platforms, data warehouses, and statistical analysis.

However, an important shift is occurring.

Organizations are increasingly discovering that many of the most important insights are not hidden within individual datasets.

They are hidden within relationships.

Knowledge Graphs and the Power of Relationships

This is where knowledge graphs enter the conversation.

Traditional databases are optimized for storing things.

Knowledge graphs are optimized for storing relationships.

Instead of asking:

What information do we have?

Knowledge graphs allow us to ask:

How are these things connected?

For a school district, those connections might include relationships among:

  • Students
  • Teachers
  • Programs
  • Community organizations
  • Postsecondary institutions
  • Employers
  • Workforce initiatives

A knowledge graph makes it possible to visualize and analyze those connections as a network rather than as a collection of disconnected spreadsheets.

Technologies such as Amazon Neptune support this type of graph-based architecture by allowing organizations to model entities and relationships as connected networks rather than simple tables.

This distinction is significant because most community challenges are not resource problems.

They are relationship problems.

Network Theory and Community Intelligence

Once relationships can be modeled, network science becomes possible.

Network analysis helps answer questions that traditional reporting systems cannot.

For example:

  • Which organizations connect otherwise disconnected groups?
  • Which partnerships create the strongest outcomes?
  • Where are critical relationships missing?
  • Which individuals or institutions act as bridges across sectors?

These questions are especially important in rural communities.

In many cases, community success depends not on the size of an organization but on its ability to connect people, resources, and opportunities across institutional boundaries.

This aligns closely with concepts such as:

  • Granovetter’s theory of weak ties
  • Burt’s structural holes
  • Network weaving
  • Collaborative governance

The most valuable institution in a community is not always the largest one.

Often it is the organization that creates connections between otherwise disconnected groups.

Knowledge graphs and network analysis help make those relationships visible.

Wisdom: What Should We Do Next?

The final stage of the journey is wisdom.

Knowledge identifies patterns.

Wisdom guides action.

A dashboard might reveal that students who participate in Career and Technical Education programs have stronger educational outcomes.

That is knowledge.

The decision to expand those programs is wisdom.

A network analysis might reveal that a specific community partnership consistently produces positive outcomes.

That is knowledge.

Choosing to strengthen that partnership is wisdom.

Technology cannot make those decisions.

People do.

This is why wisdom ultimately remains a human responsibility.

Why the Information Scientist Matters

Many people assume that data engineers are the most important professionals in a modern data ecosystem.

Data engineers are essential.

They build the infrastructure.

But infrastructure alone does not create understanding.

The information scientist serves a different role.

I often describe it this way:

The engineer builds the library.

The information scientist organizes the books.

The community decides what story they tell.

Information scientists provide the connective tissue that allows data to be transformed into knowledge and knowledge into action.

They focus on:

  • Meaning
  • Context
  • Relationships
  • Provenance
  • Ontology
  • Semantic interoperability
  • Knowledge representation

In other words, they are stewards of understanding.

Building a Learning Ecosystem

The real promise of initiatives such as MiGreatDataLake is not technological.

It is social.

The goal is not merely to centralize information.

The goal is to create a system that helps educators, researchers, policy makers, community organizations, and citizens learn together.

Data tells us what happened.

Information helps us understand it.

Knowledge explains why.

Wisdom helps us decide what to do next.

The future belongs not to the organizations with the most data, but to the communities that can transform data into shared understanding and shared understanding into collective action.