Why Your Attribution Tells You What Happened, Not Why

Marketing dashboards describe what happened, not why. HubSpot's deal source is inherited from one contact, so deal reports inherit the same flaw.

John Kelleher
John Kelleher

Pipeline from your strongest segment fell last quarter and the board wants to know why. The dashboard can show that it fell, split by channel, by month and by campaign. What it cannot show you is what caused it, so the budget conversation is settled by whoever argues most confidently rather than by the evidence.

That gap is not a charting problem and it does not close by putting AI on top of the same numbers. It is a data model problem, which makes it an engineering job.

Why the reporting describes rather than explains

Most marketing reporting is built on the contact record: contacts created, the source stamped on them at creation, and the forms they submitted. Revenue is not carried by the contact. It is carried by the deal, and behind the deal, by the company. That mismatch produces two failures, and both are mechanical.

Buying-committee dilution. A UK B2B purchase at companies from around £3m revenue upwards usually involves several people from the same organisation, arriving in your database by different routes over months. Contact-level source credits whichever one arrived first. Every other route disappears. You are not measuring what caused the deal, you are measuring which member of the committee filled in a form earliest, which is often the person with the least influence over the decision.

The long-tail account. A company that has been in your database for three years opens a deal today. Contact-level source attributes that deal to whatever brought somebody in three years ago, typically a piece of content nobody has touched since. The attribution is technically accurate and commercially useless.

Moving to deal-level reporting inherits the problem rather than solving it

The standard advice is to stop counting contacts and report on deals instead. The direction is right, but it does not escape the dilution, and that is what decides whether you have a reporting job or an engineering job.

HubSpot documents its default deal property Original Traffic Source as:

"the original traffic source of the associated contact with the earliest activity for the deal. If there's no associated contact record, the associated company's data will be used instead."

Source: HubSpot knowledge base, knowledge.hubspot.com/properties/hubspots-default-deal-properties, retrieved 16 Aug 2026.

Read that slowly. The native deal source is not determined independently at deal level. It is inherited from a single contact, the one with the earliest activity. So the committee problem and the long-tail problem do not disappear when you switch to deal reports. They follow you into them, one layer up, now carrying the authority of a revenue object. Whoever touched first takes the credit, and on a multi-stakeholder sale that is frequently somebody who did not matter.

A second constraint is worth knowing before you plan around any of this. HubSpot's knowledge base describes three attribution data sources, and two of them carry a parenthesis. Deals attribution "measures which sources, assets, and interactions had the greatest impact on generating deals (Marketing Hub Enterprise only)." Revenue attribution "measures which sources, assets, and interactions had the greatest impact on revenue (Marketing Hub Enterprise only)." Source: knowledge.hubspot.com/reports/create-attribution-reports, retrieved 16 Aug 2026.

Subscriptions change and portals differ, so check what your own subscription covers rather than taking a blog post's word for it. If you are on Professional, the honest options are an upgrade or an engineered attribution layer, and which is cheaper depends on your estate.

The question to ask instead

"Which channel produced the most leads" is answerable and close to worthless. A lead count measures friction, not demand. Remove a field from a form and it goes up, gate an asset and it goes down, and neither movement tells you anything about pipeline. It can be doubled in a fortnight by lowering the bar for what counts.

The question worth engineering for is: which channel produced revenue we actually collected. That version survives contact with a finance director, and it is harder in a way that matters, because closed-won is not collected revenue either. Deals get part-delivered, discounted at contract stage, phased across billing periods, credited back and churned. Attribution that stops at the CRM deal value points at a number that may never have arrived in the bank.

Reconciliation to the finance system is therefore part of the model, not a tidy-up afterwards. Where the finance system is Xero, Sage 200, Microsoft Business Central, NetSuite or another of the systems in our integration library, that reconciliation is a scheduled pipeline with written matching rules and an exception queue, rather than a monthly export and a spreadsheet only one person understands.

What has to be true in the data before that question is answerable

Six conditions, in order, because each depends on the ones above it.

  1. Interactions resolve to the company as well as the contact. Committee activity has to aggregate to the buying unit, otherwise deal-level reporting inherits contact-level dilution.
  2. A separate, protected, deal-level source property exists. Not the native inherited one. A value written at deal creation under rules you control, covering the cases the native property handles badly: deals created manually by a salesperson, deals created by an integration, deals for accounts that have been in the database for years, and deals where the earliest-activity contact is not the one who caused anything.
  3. Nothing downstream overwrites it. Workflows, imports, integrations and helpful humans all rewrite properties. The attribution property is write-restricted, every change to it is logged with its origin, and a backfill is a reviewed operation rather than a bulk edit run on a Friday afternoon.
  4. First touch, last touch and the full ordered interaction set are held separately. Do not collapse them into one number. A single baked-in model is why two people build the same report and get different answers. Hold the components and choose the model at report time, where the choice is visible and arguable.
  5. Offline-originated deals have their own explicit rule. Referral, outbound, partner and event deals need to be classified deliberately. Without a rule they get silently credited to whatever web session happened to occur nearby, which systematically overstates the channels that are easy to track and understates the ones that produce the revenue.
  6. The definitions live in one place. What counts as a lead, a source, a touch, an opportunity and a qualified account is written down once and referenced by every report. If the number changes depending on who built the report, you do not have one attribution model, you have several.

This is the substance of attribution and reporting engineering for marketing teams: the data model, the pipelines that feed it, and the controls that keep it trustworthy once other people start touching it.

Where AI sits in this, and where the approval line sits

AI is not the attribution model. The model is data engineering and it has to be right before anything reads from it, because a language model on top of a wrong number produces a confident wrong sentence, which is worse than the chart you started with.

What an agent does is read the assembled deal record and its interaction history and draft the explanation: what moved against the prior period, which accounts drove the movement, and which interactions preceded the deals that closed rather than the ones that did not. It separates observation from causal claim, labels inference as inference, and links every statement to the records it came from so a human can check the working.

Every action sits on a ladder before anything is built, and autonomy is granted per action, not per agent, so the same capability can read your CRM freely and still be hard-blocked from writing to it. For attribution, two lines are set as standard. Changes to the protected source property are restricted: an agent may propose a reclassification and show the effect on the historic series, and a named person commits it. Reporting definition changes are restricted for the same reason, because a definition that changes quietly is worse than a number that is visibly wrong. The steady state for the explanation layer is draft or recommend, permanently, and that is a design decision, not a limitation we are apologising for.

The honest limits

Attribution is a model. It is never a measurement, and any system that reports a single confident number for every deal is concealing its uncertainty, not resolving it.

You cannot observe causation in CRM data. You observe sequence, and sequence is not cause. Much of what moves a deal never touches a system you own: a recommendation from a former colleague, a conversation in an industry group, a podcast heard on a drive. Self-reported "how did you hear about us" answers are evidence rather than truth, and they routinely disagree with tracked source. Consent choices mean a proportion of sessions are correctly never tracked at all.

So the test of a model is not whether it is true. It is whether it is stable, documented, reproducible by two different people, and honest about what it cannot see. That is achievable, and it is enough to change how a budget decision gets made.

How you would know it worked

Not a dashboard. The measure is whether a question of the form "why did pipeline from this segment fall this quarter" gets an answer with records behind it inside the same meeting, instead of becoming an action item for a fortnight.

Alongside that, four things you can count: the proportion of deals carrying a complete, protected, deal-level source value; manual reconciliation touches per month; the variance between CRM closed-won and invoiced revenue by channel, and whether it narrows; and the correction rate when a human reviews a drafted explanation. Report pipeline and conversion indicators where the evidence supports them, but do not relabel them as revenue, and do not count saved time as cost reduction until somebody removes the cost.

The next step

Work of this kind would be one wave inside the AI Accelerator programme, a 12-month AI transformation delivered department by department, one quarter at a time, where a standard wave activates two capabilities in production rather than a catalogue. Marketing is the third department by default and the order can be changed.

If you want to know whether your data can support a causal question today, the next step is the AI and Data Readiness Assessment, one of five short fixed-scope assessments on our diagnostics page. It looks at your CRM, the systems feeding it and the reports you run now, and comes back with three things: whether deal-level attribution is achievable on your current data, what your own subscription already covers so you do not pay anybody to rebuild it, and what the first piece of engineering would be. Fixed scope, no obligation. Attribution normally lands in the third quarter of a programme, for reasons explained in how to sequence an AI rollout.

John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.