21 June 2024
Tracking data platform usage is critical
Your team has invested significant time and money into building data infrastructure and creating an array of data assets - raw data tables, transformed data sets, dashboards, machine learning models, and more. But has this activity truly generated value?
Your team has invested significant time and money into building a robust data infrastructure and more broadly, the organization has created a wide array of data assets - raw data tables, transformed data sets, dashboards, machine learning models, and more. But has this activity truly generated value?
Too often, we see teams taking a "build it and they will come" approach. They make the upfront investment of developing data pipelines, data models and dashboards, along with interfaces to discover and interact with these assets, but then lose focus on how any of this is being used. Whether through lack of attention or technical difficulties, this can lead to costly inefficiencies and the data team being viewed as a cost center rather than a source of value. Just as you wouldn't run your business without an eye on customer numbers, CSAT, revenue and costs, monitoring data platform usage is important and brings significant benefits:
Allocate human effort efficiently
When you have visibility into which dashboards, visualizations or models are heavily utilized, you can make smarter decisions about which data tables and pipelines to monitor more closely. Data trust can be prioritized based on the way in which assets are being used, prioritizing completeness over accuracy for example. Remediation of issues can be prioritized based on customer impact. Data teams are often chronically stretched, so every bit of efficiency and prioritization here helps.
Match your data assets with needs
Looking at usage patterns indicates which assets are in high demand across teams/functions. This allows owners to engage with their consumers and strategically build out the assets that will have the biggest impact. When assets are reviewed and improved, overall information quality increases, duplication reduces and insight generation is far easier.
Improve security and management
Understanding usage helps with data governance. Sensitive data can be reviewed to make sure that it is only being accessed on a need-to-know basis. Sanitization may be possible on popular content which commonly triggers access controls. Asset ownership can be re-allocated as employees join or leave the company. When summarized by team, it provides a great window into where data-driven decisions are being made across the company. Better data management reduces risk and improves insight.
Associate cost with usage
Data asset usage extent and depth isn't the entire picture of value generated from your data assets, but it is a great first step (Aside: Keep an eye out on our blog for more content on this topic!). Immediately, underutilized assets can be deprecated or archived, and accessibility of heavily used assets can be improved. When combined with costs, it allows for a glimpse of that holy grail, determining return on data investment. Doing so is a huge benefit for data platform management, as it allows clear-eyed decisions to be made about where to focus.
Optimize user experience
So much of the data user experience goes beyond the count of queries on a table in the data warehouse. There is a full user journey from business problem to data asset discovery to analysis and insight, much of which doesn't involve writing a SQL query on one particular table. Understanding the end to end user experience provides great insight into what features to build into your data environment - prioritizing discoverability, access, understanding, trust or utilization.
The hard part: Making it a reality
The benefits of tracking data asset usage should now be abundantly clear. When turning to how to make this a reality, the picture becomes much less straightforward.
Many SaaS vendors don't offer any insight into this type of information, and particularly not at the initial tier of their pricing offerings. The data user journey cuts across a variety of internally-developed and SaaS tools. Some are UI-first, some are code-first. Stitching together usage intelligence across your data warehouse, BI tools, MLOps platform, reverse ETL solution, and more is a significant undertaking. As a result, if you want to go down this route, it is important to select tools that will grow gracefully with you, allowing incremental builds and ultimately providing you the right information to make the right investment decisions at the right time.
As an example, lets zero in on some popular tools in the business intelligence layer to begin with:
| BI Tool | Log availability | Lookback window | Access interface |
|---|---|---|---|
| Microsoft Excel | File edits and views; Users performing activity | 6-12 months | Microsoft Purview UI |
| Power BI | Content views; Query origins and sources; Users performing activity | 7-90 days depending on interface | Microsoft Purview UI; Logs API |
| Looker Studio | Query origins, sources and definitions; Users performing activity | 6 months | Google Workspace log explorer |
| Looker | Content edits, views and run times; Data model changes and field usage; Error details; Query origins, sources, definitions and run stages; Users performing activity | 90 days | Looker system activity dashboards, explores and API |
| Tableau | Content edits and views; Data model changes; Query origins and sources; Users performing activity; NB: Log structure differs between Tableau Cloud and Server | 30 days - 6 months | Tableau Server PostgreSQL DB; Tableau Advanced Management for Tableau Cloud, AWS S3 |
While differentiation between BI tools on data access interfaces does exist, most are similar at the core. When looking at the utility of and access to user activity and content logging, however, there is significant differentiation between the various offerings, and this is not visible up-front on vendors' marketing websites. In many cases, there is limited look-through to the underlying tables being queried in the data warehouse, cache policy effectiveness and other data to drive performance optimizations. As a result, most organizations will eventually use a combination of logs from the data warehouse and BI tools to monitor data asset usage.
Advantages of Using BI tool usage tracking:
- Out-of-the-box capabilities tailored for the BI platform
- Captures dashboard/report specific usage metrics easily
- Can track operations beyond just queries (exports, subscriptions, alerts, etc.)
- UIs typically make it easier for non-technical users to access usage data
Advantages of Using Data Warehouse Logs:
- Get full visibility into all queries, not just BI tool usage
- Can track usage of raw data sources, models, transformations, etc.
- Logs contain more granular information (query text, timing, user, etc.)
- Query logs are the true source of truth for all data access
Most comprehensive usage tracking strategies leverage both approaches synergistically:
- Use BI tool usage data for high-level dashboard/report metrics
- Parse query logs to track usage of underlying data assets
- Correlate between the two data sources for an end-to-end view
- Integrate logs into BI tool to build usage dashboards/reporting
The main TL;DR from this? If you want to drive your data team like a product team--with a focus on understanding the data user experience, building features that users want, and running an efficient data platform, you need to be collecting data from your data platform interfaces that helps to drive those decisions. Logs from your BI tooling, query engine and/or database will likely be a core part of this, so setting up data pipelines to ingest these logs will help to build a fact base over time. Our quick starts can help to reduce time-to-value here, so drop some time in for a chat if this resonates.
Cover image by rawpixel.com on Freepik