

Everyone has a love/hate relationship with attribution. Mostly hate. I was the CMO of Branch and I built a multi-touch model there that was directionally useful, not truly right. It over-credited what was easy to measure (late-stage clicks, form fills and email) and under-credited what actually created the deal (content, referrals, events, the dark funnel). For the past two years at Upside we’ve been building the data layer to fix that, and a few months ago Alex Bauer (Branch early employee, now my co-founder at Upside) built something on top of it that I think is truly new: an attribution engine that runs a panel of independent LLM analysts on every deal, argues with itself, and produces a zero-sum credit allocation where every dollar is accounted for.
This post is the technical overview of how it works. The part that most marketers will be surprised by, but eventually agree with: demo request forms (and all form fills and customer actions) get exactly 0% credit. I’ll get to why.
One thing I have learned while working on attribution at Upside and at Branch is that there is no attribution model that can be applied to every problem. The tagline “all models are wrong, but some are useful” applies to attribution models, and most of what this article explores is the concept of applying a zero-sum credit allocation to the activities that led to sourcing a complex B2B deal.
The market currently agrees that there is no accurate model to solve this and instead uses time-based heuristics: first-touch, last-touch, U-shape, W-shape. These are just weighting curves over timestamps. The first or last touchpoint might not matter at all, but the curve doesn’t know that, because the curve doesn’t read the emails.
Our thesis is that finding a more accurate B2B attribution model failed for two reasons, and neither of them is “attribution is conceptually impossible”:
The first version of an attribution model was a more complex weighted model — give weights to all the touchpoints involved based on rules: emails take less effort and should get less weight than webinars.
To know if this kind of model was right we decided to manually look at hundreds of complex customer deals ($1M+) and understand what actually drove them. We found two things:
In most cases the real story of the deal was completely different from what an attribution model showed, because any model would only look at structured data .
So before any of the interesting agent stuff, we spent about two years building the boring part: pulling CRM, email, calendar, call transcripts, and web activity into one place, then healing it. Identity resolution (the same person with two emails across two systems becomes one person), buying group detection, activity deduplication across systems, and extracting touchpoints that only exist inside unstructured text. The output is an account-level timeline that is the cleanest, most complete version of “what happened” we can reconstruct.
That gets you a great audit console but does not tell you what worked . Some of our customers’ deals run 18+ months with thousands and even tens of thousands of touchpoints. Knowing everything that happened just makes the credit question harder.
Alex’s insight was that we already had an intelligence layer that agents use to query this data , and now we could use AI to replicate the human logic we did in the early days. So instead of fitting a model, what if we made AI agents do what a smart human analyst does: read the whole deal, reason about causality, and defend a credit allocation with citations?
That required building a real domain theory first, because “just ask the LLM” produces vibes, not attribution. The theory lives in two prompt documents every agent reads: a commander’s intent (what we’re building, for whom, and the governing principles) and structural rules (the mechanics). Our main concepts:
Every discrete action on a deal timeline is one of three things:
The rule that everything else is based on, which can be quite controversial: credit belongs to whoever instantiated the thing, not to the buyer’s observable reaction to it. Responses always carry 0% credit. They are evidence that something worked, not the thing that worked.
This is why the demo request gets zero. From the buyer’s perspective, the demo request form is the annoying thing they have to get through to talk to someone about buying. It did nothing to convince them. If anything it was friction. Instead, everything that led them to be convinced enough to fill it out is what deserves the credit. Most attribution models we’ve seen, including the one I built at Branch, hand form fills 20% of the deal. Then leadership reads the report and concludes the growth strategy is “more forms.”
And if we can’t see anything before the demo request? What happened after doesn’t get more credit — instead the pool of circumstantial influence gets the credit, more on that below.
The 0% rule is also a structural constraint, not just philosophy: in our classification hierarchy, bids roll up through the selling org branch (team → person → channel → campaign → engagement) and cobids through the third-party branch (entity type → actor). A response with non-zero credit belongs to neither branch and shows up as “Unassigned” in every rollup. Making responses 0% keeps the whole reporting tree zero-sum and aggregatable.
Touchpoints come in three evidence classes:
Direct and extracted touchpoints each get a specific credit percentage. The remainder goes into a pool of circumstantial influence, broken into named, explicitly probabilistic subdivisions (organizational readiness, internal champion relay, brand awareness, and so on). The pool is not a hedge, but a specific claim about invisible influence, and its size is computed from signal coverage:
A deal with a 10% pool has great CRM coverage. A deal with a 65% pool is telling you that most of your pipeline creation is invisible to your systems, which is itself one of the most useful things a CMO can learn.
A run starts with “run attribution on the Acme deal.” An orchestrator dispatches three independent analyst agents . Each is an Opus-tier model with the same CRM query tools and the same instructions, running up to an hour. They are not assigned different angles on purpose. The point is that three equally capable analysts looking at the same messy deal will not reach identical conclusions, same as three skilled humans wouldn’t. Each one’s primary obligation is traceability: every claim needs citations (record IDs, transcript quotes, email references) a downstream agent can verify.
Their outputs go to a consensus judge . The judge weighs the reasoning, not the votes. A year ago Alex and I were doing this manually, disagreeing about deals and arguing it out, and the judge behaves the way we did: where the runs agree, follow. Where they disagree, evaluate why . A minority run that noticed evidence the others missed can override the majority, but overriding requires explicit reasoning. We deliberately rejected majority vote and median-of-three, because the entire value of the system is contextual reasoning, and mechanical aggregation throws that away. When the spread is too wide to adjudicate, the judge escalates: 3 analysts → 5 → 7.
(Why start at three? Cost. Each analyst is an expensive model running up to an hour. On a straightforward deal, five analysts agree the same way three do, just slower and pricier.)
Then two more agents that exist mostly to be skeptical:
One prompt-design principle runs through all of them, straight from the structural rules doc: exhaust before escaping. Every agent has fallbacks (a classification path can stop at team level, a source can be “unknown”) but using one requires documenting what you tried first. A short path means you looked everywhere and truly could not resolve deeper, not that it was hard and you moved on. Agent time is cheap; a misleading value in a CMO’s report is not. Relatedly, we made classification paths variable-depth with explicit null gaps, because “we know the team and channel but not the person” is a more trustworthy statement than a fabricated person name, and downstream consumers can see exactly where knowledge stops.
Per deal, two outputs: a zero-sum multi-touch allocation across all identified influences, and an opportunity source (the system distinguishes the hand-raise , the buyer’s decision moment, from the source , the upstream event whose removal would have broken the causal chain). Here’s a real (anonymized) deal:
Note the RFP email flagged HAND-RAISE with $0 credit, while the AE’s relationship maintenance during dormancy, the SDR’s 12-month outbound sequence, and the webinar carry the dollars. And every touchpoint, including the $0 ones, carries a reasoning trace exposed in the UI, so when someone asks “why did the chatbot get nothing but the chatbot campaign get credit,” the answer is right there.
The causal chain reads like a deal post-mortem written by someone with perfect memory: persistent SDR outreach → webinar attendance → post-webinar follow-up → discovery and demos → budget dormancy → relationship maintenance → buyer re-initiates RFP. Then everything rolls up across the portfolio by team, channel, campaign, and salesperson, which is where it starts looking like a normal attribution dashboard, except the numbers underneath are built from evidence instead of curve positions.
For the times when leadership does want the ONE thing that got the deal, our model also allows you to see the one activity that got this deal moving, and run aggregate reporting using that model as well.
The interesting part is that this approach finds sources that were thought impossible to detect: “there is no way your model can tell this deal came through an alumni network.” It could. Someone had mentioned it on a call, the analyst extracted it, and it survived the judge.
While we believe this approach outputs the most accurate results in zero-sum B2B attribution, we also know it has its limitations. This is one way to approach attribution powered by a clean and unified data layer, but the possibilities in what you can build to measure are endless. Examples of other types of measurement our customers built on top of Upside data can be found in our mini-app library .
We packaged the whole orchestration layer as a product called Pipedash , and the underlying data layer is queryable if you’d rather build your own version of this on top. Alex and I did a full webinar walkthrough with live examples if you want the video version. And yes, if you submit a demo request after reading this, the form itself will get 0% credit in our system. This blog post will get it instead.
Book a demo and see what your team can build, automate, and finally answer once your GTM data is something you trust.
Hacker News
news.ycombinator.com