A custom attribution model is a rule- or algorithm-driven allocation you build when off-the-shelf models misrepresent your conversion paths. You need one when your data is complete enough to trust, your channel mix is too complex for first-touch or last-touch logic, and the budget decisions riding on attribution are big enough to justify the engineering. The trade-off is real: better signal in exchange for build time, maintenance, and new privacy constraints on browser-level data.
TL;DR:
- Begin with rule based changes, such as reducing branded search credit; adopt Markov or Shapley methods when data volume and engineering capacity allow.
- Run the new model beside the current one for at least one full attribution window, then compare recommendations with holdout results before changing budgets.
- Browser privacy changes can replace individual event paths with aggregated, noisy reports and tighter lookback windows, so combine them with server side and CRM data.
- Build internally with simple channels and data in one or two connected tools; seek technical help when CRM, offline, and advertising data remain disconnected.
Table of Contents
- What a custom attribution model is and the types you can start from
- Why build a custom attribution model (and where it goes wrong)
- Available customizations and example credit rules
- How to build and deploy a custom attribution model step by step
- Algorithmic approaches: Markov chains, Shapley values, and counterfactual Shapley
- Privacy, browser APIs, and what you can still measure
- Implementation checklist and validation tests
- How a technical partner approaches a custom attribution build
- Deciding between DIY and a technical partner
- Where Dogtooth Inc fits in a custom attribution project
- FAQ
- Sources
What a custom attribution model is and the types you can start from
Every attribution project starts with a baseline. Linear splits credit evenly across every touchpoint. Time decay weights recent touches more heavily, and position-based models (often a 40/20/40 or similar split) favor the first and last interaction while giving a smaller share to the middle.
A custom model takes one of these as a starting point and changes the parts that do not fit your actual customer journey. That usually means adjusting the lookback window, rewriting the weighting formula, adding credit rules for specific channel behaviors, or excluding touchpoints that do not represent real marketing influence, like internal redirects or support visits.
Customizations split into two broad categories:
- Rule-based customizations: manual adjustments like reducing credit for branded search, excluding bot traffic, or capping how much weight a single channel can absorb.
- Data-driven algorithmic methods: statistical models, such as Markov chains or Shapley values, that calculate credit from the actual structure of your conversion paths rather than a fixed formula.
Vendor documentation from Amplitude’s attribution framework guide maps these categories directly to the configuration options most analytics platforms expose, which makes it a useful reference when you are deciding where to start. Most teams begin with rule-based logic because it is faster to explain to stakeholders, then move toward algorithmic methods once they have the data volume and engineering support to justify it.
Why build a custom attribution model (and where it goes wrong)
The upside of a custom model is straightforward: it lets you allocate budget based on how channels actually influence conversions in your specific funnel, not a generic assumption baked into a reporting tool. A B2B company with a six-month sales cycle and a direct-to-consumer brand with a same-day purchase cycle should never be using the same attribution logic, and a custom model is how you fix that mismatch.
The pitfalls tend to cluster in three places:
- Data gaps: missing offline conversions, broken cross-device stitching, or untracked dark social traffic that quietly understates certain channels.
- Overfitting: building a model so tailored to last quarter’s data that it breaks the moment buyer behavior shifts.
- Bad assumptions: treating a correlation (a channel that shows up a lot) as causation (a channel that actually drives conversions).
A review of heuristic versus algorithmic approaches in marketing measurement research found that heuristic models produce inconsistent results once journeys get long and multichannel, which is exactly the scenario most teams are trying to solve with a custom build in the first place.
The practical fix for most of this is validation before rollout: run the new model alongside your existing one for a full reporting cycle, and compare budget recommendations before you let either model drive spending decisions alone.
Pro Tip: Run your custom model and your old model in parallel for at least one full attribution window before retiring the old one, so you can see exactly where and why the numbers diverge.
Available customizations and example credit rules
Most attribution platforms expose a similar set of knobs once you move past the default model:
- Lookback window: how far back a touchpoint can still receive credit, often 30, 60, or 90 days.
- Half-life: how fast a time-decay model’s credit shrinks as a touchpoint ages.
- Impression vs. click multipliers: giving clicks more weight than passive impressions, or the reverse for upper-funnel channels.
- Position splits: the percentage assigned to first touch, last touch, and the middle interactions.
- Include and exclude rules: filtering out internal traffic, support sessions, or specific referral domains that do not represent marketing influence.
From there, teams typically write a short set of custom credit rules on top of the base model. A few common examples: reduce credit for branded search terms since those clicks often represent demand you already created, boost credit for any touchpoint that occurs within two minutes of a paid click because it signals an active, high-intent session, and assign zero credit to internal redirects that exist for technical reasons rather than persuasion.
Here is how one of those rules changes outcomes for a simple four-touch path, moving from a linear model to a custom model that reduces branded search credit and boosts the paid-social touch that preceded a conversion within the two-minute window:
The shift looks small on paper, but across thousands of paths it can move a meaningful share of attributed revenue away from branded search and toward the paid social channel that was actually doing the persuading.
How to build and deploy a custom attribution model step by step
Building a custom model is a sequencing problem as much as a technical one. Skip a step and you end up validating a model against data it was never designed to handle.
- Define objectives and success metrics first. Decide what question the model needs to answer, like “which channel deserves more budget” or “how long is our actual conversion window,” before touching any configuration.
- Inventory your tracking. List every data source touching the funnel: web analytics, CRM, call tracking, offline point-of-sale, and any QR-code or print campaigns that drive customers online, since gaps here undermine everything downstream.
- Centralize the data. Pull everything into one warehouse or data layer with consistent identifiers, so a single customer’s path can be reconstructed across channels and devices.
- Decide sessionization and lookback rules. Pick a lookback window that matches your actual sales cycle, not a platform default, and align session boundaries across tools so touchpoints are not double-counted or dropped.
- Implement the credit logic. Build the rule-based adjustments or algorithmic model in your chosen tool or custom pipeline, documenting every rule so it can be audited later.
- Run validation tests. Compare outputs against the previous model and against known outcomes, like a channel you deliberately paused, before trusting the new numbers.
- Deploy with monitoring. Launch with dashboards tracking model drift and data completeness, and set a review cadence, often monthly or quarterly, to catch when buyer behavior shifts enough to require retuning.
Offline-to-online measurement deserves specific attention in step two. QR codes on print, packaging, or in-store signage generate real touchpoints that are easy to lose if they are not tracked with the same rigor as digital clicks; a QR analytics guide from QRlytics walks through capturing scan-to-conversion data in a way that plugs cleanly into a broader attribution pipeline.
Pro Tip: Build your validation tests before you finish building the model itself. Deciding what “correct” looks like after the fact almost always biases the result toward whatever the new model already says.
Algorithmic approaches: Markov chains, Shapley values, and counterfactual Shapley
Heuristic models assign credit by a fixed formula. Algorithmic models calculate it from the actual structure of your data, which is why they tend to perform better once your conversion paths get complicated.
Markov chain models treat each touchpoint as a state in a customer journey and calculate a “removal effect”: how much the conversion rate drops when a given channel is removed from all paths entirely. A channel with a large removal effect gets more credit, because the model can show its absence actually hurts conversions, not just that it appears often.
Shapley value attribution borrows from cooperative game theory, calculating each channel’s fair share of credit based on its average marginal contribution across every possible combination of touchpoints. The challenge is that a naive, uniform Shapley calculation can still misrepresent contribution in advertising contexts. Research presented at The Web Conference introduces a counterfactual-adjusted Shapley metric that corrects those flaws in uniform heuristics while remaining computationally implementable in production pipelines, which makes it one of the more theoretically sound options available to teams with the engineering capacity to run it.
- Markov chains work well when you need a clear, defensible “what if we cut this channel” answer for budget conversations.
- Shapley and counterfactual-adjusted Shapley work well when you need a fairness-grounded credit split that holds up across many different path combinations.
- Higher-order Markov models capture sequence dependencies that first-order Markov misses, which matters when the order of touchpoints, not just their presence, affects outcomes.
Empirical comparisons consistently show Markov and Shapley-based methods outperforming simple heuristics on real conversion data, according to research comparing classification and Markov attribution models, though both approaches demand larger sample sizes and more compute than a rule-based model, so smaller advertisers often get more value starting with heuristics and graduating to algorithmic methods as volume grows.
Privacy, browser APIs, and what you can still measure
Browser-level privacy changes have reshaped what granular attribution data is even available to collect, and any custom model built today has to account for that from the start.
The W3C’s Private Attribution specification defines a browser API design that returns aggregated conversion reports under differential privacy, with core parameters including a privacy budget (epsilon), histogram buckets, and a capped lookback window rather than raw, individually identifiable event logs. Google’s parallel implementation, detailed in its Attribution Reporting summary reports documentation, batches conversion data into aggregatable reports processed through a trusted execution environment, which changes the shape of the data you receive, not just its source.
The practical effect is less granularity and more noise, especially for multi-touch paths where you once could see every individual step.
- Expect coarser data: aggregated buckets instead of per-user event streams.
- Expect added noise from differential privacy, which grows as you slice the data into smaller segments.
- Expect tighter lookback windows that may not match your actual sales cycle.
Representing impressions as nodes rather than full paths reduces aggregation noise and preserves more stable signals when building multi-touch reports under these constraints.
That guidance, from Privacy Sandbox’s private aggregation documentation, points to the main mitigation strategy available right now: build hybrid measurement that blends aggregated browser reports with server-side and CRM data, treat any single noisy aggregate as a trend indicator rather than a precise number, and set conservative decision thresholds so a budget shift requires a clearer signal than it would have under old, cookie-based tracking.
Implementation checklist and validation tests
Before launch, run through a short preflight checklist:
- Confirm a consistent schema across every data source feeding the model.
- Check for completeness, especially offline and CRM touchpoints that are easy to under-track.
- Remove duplicate events that inflate a channel’s apparent touchpoint count.
- Verify consistent identifiers so a single customer’s path stitches together correctly across devices.
Once the model is live, a small set of validation experiments catches most problems early:
- Run a holdout group that excludes one channel entirely and compare the model’s predicted impact against the real drop in conversions.
- Use channel removal tests to sanity-check that the model’s credit allocation lines up with observed cause and effect.
- Run budget-shift checks: move a small, controlled amount of spend based on the model’s recommendation and measure the actual outcome before committing the full budget.
For ongoing monitoring, track data completeness rate, model drift between reporting periods, and the variance between predicted and actual conversion lift. A practical guide to ROI tracking covers monitoring setups that pair well with a custom attribution rollout once the model itself is validated.
How a technical partner approaches a custom attribution build
A typical engagement runs through five phases: discovery (mapping existing data sources and gaps), pipeline build (centralizing tracking into one consistent layer), model configuration (choosing rule-based or algorithmic logic and tuning it to the client’s actual sales cycle), QA (running the validation tests above against real data), and iteration (a set cadence for retuning as channel mix shifts).

We build AI agents that work inside a client’s actual workflow rather than as a bolt-on reporting layer. The same agent infrastructure that drafts emails, checks compliance, and books sessions for other projects can run scheduled monitoring checks and flag attribution drift automatically, instead of waiting for a monthly manual review.
A client is usually ready to bring in a technical partner once internal data is scattered across tools nobody fully owns, or once attribution decisions are shaping a budget large enough that a wrong model is expensive.
— Chase Weir
Deciding between DIY and a technical partner
Build it yourself when your channel mix is simple, your data already lives in one or two connected tools, and the decisions riding on attribution are modest. Bring in a partner once you are juggling CRM, offline, and multi-platform ad data with no clean way to join it, or once the budget decisions depend on getting the model right the first time.
A practical path: start with a small pilot focused on one funnel or product line, validate the approach against real outcomes, then move into a retained optimization arrangement once the model proves its value and needs ongoing tuning as channels and buyer behavior shift.
Where Dogtooth Inc fits in a custom attribution project
We build the pipeline work that custom attribution depends on: centralizing scattered tracking data, building the platforms and dashboards that surface model output, and designing AI agents that run monitoring and validation checks without a manual review cycle every time. An email deliverability platform can provide high-volume senders direct measurement of email performance instead of relying on third-party tracking that privacy changes have made less reliable.

The outcome we aim for is a repeatable pipeline you can keep tuning, not a one-time report. If your attribution project depends on data engineering, custom dashboards, or workflow automation more than off-the-shelf software, reach out through our services page to talk through a pilot.
FAQ
Which attribution model is best?
There is no single best model; the right choice depends on your conversion path length, data completeness, and engineering capacity. Simple funnels often do fine with heuristic models like linear or position-based, while longer, multichannel journeys tend to benefit from algorithmic methods like Markov chains or Shapley value, which empirical comparisons show outperforming heuristics as complexity grows.
What does attribution model mean?
An attribution model is the method used to assign credit for a conversion across the different marketing touchpoints a customer interacted with before converting. It can be as simple as crediting one touchpoint entirely, as in first-touch or last-touch, or as complex as a statistical model that calculates each touchpoint’s contribution from the structure of thousands of customer paths.
What are the different types of attribution models?
The common baseline types are first-touch, last-touch, linear, time decay, and position-based models, each of which splits credit across touchpoints differently. Beyond these heuristics, algorithmic types like Markov chain and Shapley value attribution calculate credit from actual path data rather than a fixed rule, and vendor documentation maps all of these to configuration options available in most analytics platforms.
What are examples of custom attribution rules?
Common custom rules include reducing credit for branded search since that demand was often already created by other marketing, boosting credit for any touchpoint occurring within a short window before conversion to flag high-intent sessions, and excluding internal redirects or support visits that carry no real persuasive influence. Teams typically layer two or three of these rules on top of a baseline model like linear or position-based rather than replacing the whole framework.



