Blog home

Building your own marketing reporting tool is getting easier. Trusting the data isn't.

05 Oct 2026•8 min read•Author: Nick Beno

Over the last couple of months, we've had several conversations with marketing agencies that are building their own reporting and analytics tools in-house.

And I think that's a really positive development.

A couple of years ago, much of the conversation around AI in agencies was defensive. What would it automate? Which services would become commoditised? What would clients start doing themselves?

Now we increasingly seeing the opposite.

Agencies are experimenting.

They're connecting marketing platforms, building internal applications, automating parts of reporting and creating AI tools that analyse performance or produce client-facing outputs.

A fairly common setup might involve a connector platform bringing the data in, a BI tool such as Looker Studio presenting it, and an AI layer such as Claude performing a specific task on top.

The barrier to building something useful has fallen dramatically.

But one part of this conversation isn't discussed nearly enough.

Building the thing is increasingly becoming the easy part. Maintaining trustworthy data underneath it is not.

And GA4 is one of the best examples of why.

For many marketing teams, GA4 has become one of the central measurement layers connecting acquisition activity with what actually happens on a website.

Someone clicks an ad.

They land on a page.

They browse.

They convert.

They purchase.

They return later through another channel.

GA4 often sits somewhere in the middle of the questions marketers are trying to answer about that journey.

Which makes it incredibly valuable.

It also means that when agencies start building their own reporting infrastructure, GA4 frequently becomes one of the most important datasets they need to work with.

At first, that can seem straightforward.

Connect GA4.

Pull the required metrics.

Put them in a dashboard.

Add Google Ads, Meta, HubSpot or whichever other platforms matter.

Perhaps add an AI layer to analyse the results.

You now have something that would have taken far more engineering resource to build a few years ago.

But the first version is rarely the difficult part.

What happens next is.

Imagine you've built an internal dashboard for a client.

Initially they want Sessions by Channel.

No problem.

A few weeks later they ask for Performance by Device.

Then Channel by Device.

Then Monthly Active Users.

Then Landing-page Performance.

Then Ecommerce data.

Then Campaign Reporting.

Then a custom GA4 dimension gets added.

Then somebody asks why the number in the dashboard doesn't match what they're seeing inside GA4.

None of these are unreasonable requests.

In fact, this is exactly what should happen when reporting becomes useful.

People start asking better questions.

But every additional question creates new requirements further down the reporting stack.

And that's where the difference between connecting data and modelling data starts to matter.

This is one of the most important lessons our engineer, Ehsan Honarbakhsh, surfaced while working through GA4 data modelling.

GA4 doesn't behave like a conventional warehouse table where you collect a giant dataset once and then rearrange the rows however you like.

The question you ask GA4 matters.

If you request Sessions by Channel, GA4 calculates the result for that particular breakdown.

If you request Sessions by Device, it calculates that breakdown.

Having both datasets doesn't automatically mean you can accurately reconstruct sessions by Channel and Device afterwards.

If you need that combination, you may need to design for that combination.

That sounds like a small technical detail.

Strategically, it's much more important.

Because it means that as reporting requirements evolve, the underlying architecture often has to evolve with them.

The reporting layer isn't a finished asset sitting on top of static data.

It is a living system.

This is where the falling barrier to software development can create false confidence.

Modern tools are extremely good at producing interfaces.

You can connect an API.

Ask an AI coding tool to build a chart.

Add a filter.

Create a table.

Generate some commentary.

And visually, everything can look fantastic.

The dangerous problems are the ones that don't produce an error message.

GA4 contains plenty of situations where a perfectly valid query can return numbers that become misleading when used in the wrong way.

Users, for example, aren't necessarily additive across different dimensions or periods.

A person may appear across multiple days, devices or channels.

Adding those rows together doesn't create a unique-user total.

Sessions create their own challenges depending on the dimensions they're analysed against.

Even the grain of the report matters. A number requested directly for a month shouldn't automatically be assumed to equal the total you get by adding together every daily row within that month.

There is nothing visibly broken about the dashboard.

There is no giant red warning saying:

DO NOT USE THIS NUMBER.

You simply get a number.

And that is exactly why the quality of the underlying data model matters so much.

  1. Connection — getting data out of the platforms.
  2. Modelling — structuring that data so it answers the questions you actually need to ask.
  3. Validation — ensuring those answers reconcile and behave as expected.
  4. Presentation — turning the data into dashboards, charts and reports.
  5. Analysis — interpreting what happened and deciding what to do next.

Over the last few years, layers one, four and five have become dramatically more accessible.

Connector platforms have made it easier to move data.

BI platforms have made visualisation accessible to almost everyone.

AI has made analysis, querying and even software development far easier.

But layers two and three haven't disappeared.

If anything, they become more important as everything around them gets easier.

Because once you can build ten new reports in the time it previously took to build one, you can also create ten new ways for the underlying data to be misunderstood.

The interface is moving faster.

The data architecture still has to keep up.

Traditionally, companies dealing with this level of complexity would have people whose job was specifically to worry about it.

Data engineers.

Analytics engineers.

Data analysts.

People who understand the grain of datasets, metric definitions, attribution logic, API limitations and how different tables should be used.

The challenge for many marketing agencies is obvious.

Most don't have the economics to maintain an entire internal data team.

An agency might have brilliant paid media specialists, SEO specialists, strategists and account managers.

Hiring engineers to continually maintain reporting architecture across dozens of clients is a different proposition entirely.

This is why agencies need to be careful when calculating the cost of building internally.

The question isn't simply:

How quickly can we build the first version?

It is also:

Who owns it afterwards?

Who checks the data when the API changes?

Who investigates discrepancies?

Who understands why two legitimate GA4 queries return different numbers?

Who updates the model when a client wants a breakdown nobody predicted six months earlier?

Who tests that a new feature hasn't quietly changed something that already worked?

That is where the real cost starts to appear.

Quite the opposite.

I think agencies should be experimenting aggressively right now.

There is a huge opportunity to build software around the way individual agencies actually work.

Every agency develops its own methodologies, reporting processes and ways of analysing accounts.

For the first time, building bespoke technology around those processes is becoming accessible to organisations that would never previously have considered themselves software companies.

That's exciting.

But there is an important distinction between experimentation and infrastructure.

You can vibe-code a reporting interface over a weekend.

If that interface starts informing client strategy on Monday morning, the standard of engineering underneath it changes considerably.

The more consequential the decision, the more trustworthy the data feeding that decision needs to be.

A dashboard can have beautiful visualisations.

An AI system can write an impressive analysis.

A report can automatically identify that conversion rate has fallen 14%.

But the first time somebody in a client meeting asks:

"Why does this number say 42,000 when GA4 says 39,500?"

everything changes.

Because once somebody starts doubting one number, they often start doubting the rest.

And once trust in reporting disappears, the quality of the dashboard almost stops mattering.

This is something we've thought about a lot while building PolyBox.

The visible reporting layer naturally gets a lot of attention. That's what people interact with.

But much of the important work happens underneath it: ensuring connectors, data models and reporting logic can support the different ways agencies need to analyse performance.

Particularly with GA4, one of our priorities has been making sure the data can be presented in ways that reconcile with how agencies expect to see it within GA4 itself.

Because clean data creates trust.

And trust is what allows the conversation to move beyond reporting towards the thing that actually matters:

What should we do next?

We also know that not every agency wants to use a reporting platform.

Some teams genuinely want to build their own infrastructure.

And if that's you, we'd rather make some of what we've learned useful.

Our engineer Ehsan Honarbakhsh has put together aguide to GA4 data modelling, covering the principles behind building GA4 Data API tables that behave correctly, including scope rules, additivity, table design, BigQuery models and validation.

You can find the full guide here:

https://github.com/ehsan-honarbakhsh/GA4-Data-Modeling

It goes considerably deeper technically than we have here, deliberately.

Because the wider point isn't that every agency founder needs to become a data engineer.

It's that somebody, or something, in your reporting architecture needs to take responsibility for the problems data engineers have historically solved.

AI has dramatically lowered the barrier to creating software.

That is going to lead to some genuinely brilliant tools being built inside marketing agencies over the next few years.

But the dashboard is only the part everyone sees.

The data model underneath it is the part that determines whether anyone should trust it.