I started with a membership question. I ended up separating three things that are easy to confuse: who joins, what they buy, and what they actually pay.
Imagine a retailer offering discounts and reward coins through a paid membership. Members spend more than everyone else. That sounds like a success, until you ask whether they would have spent more anyway.
I built this as an independent simulation with 40,000 synthetic customers and 573,352 transactions. No real customer records or company engagement are involved. The point was to build an analysis I could test against a known simulated counterfactual.
After launch, members spent 22.3% more per day than nonmembers. But when I looked backward, eventual members were already spending 40.6% more before membership existed.
That changed the question. Comparing the two groups after launch was mixing the program with the kinds of people who chose it. The pre-launch gap is evidence of selection, not an estimate of the program effect.
40.6% higher spending among eventual members.
Unmatched group averages. Switching periods does not estimate a causal effect: the groups already differ before launch.
Source and assumptions ↗I matched adopters to non-adopters with similar observed characteristics, then compared how spending changed in the two groups. Matching and difference-in-differences are stages of one design here, not two independent confirmations.
The final analysis retained 12,153 pairs. It still depends on assumptions, especially that their spending would have followed parallel trends without the program. Similar starting points help, but they cannot prove that assumption.
The estimated change in product value before discounts was +$0.1367 per customer per day. The estimated change in net product payments was −$0.0103/day, with a 95% interval from −$0.0232 to +$0.0027.
Those results tell a more useful story than “members spend more.” In this simulation, the contrast is consistent with purchasing responses being offset by discounts and redeemed rewards. It does not establish higher net payments, and an interval crossing zero does not prove the effect is exactly zero.
Net product payments exclude membership fees. Product value is measured at modeled pre-discount prices, not in physical units. Neither outcome is profit; costs and revenue recognition need their own analysis.
| Outcome | Estimate | 95% interval |
|---|---|---|
| Product value | +$0.1367 | +$0.1226 to +$0.1509 |
| Net payments | −$0.0103 | −$0.0232 to +$0.0027 |
Dots are matched DiD estimates; lines are 95% intervals. The 457-day view only rescales the same estimate, not a new forecast. Product value uses pre-discount prices; payments exclude membership fees. Neither measures profit.
Source and assumptions ↗One pretrend diagnostic raised a warning: raw p = 0.0165, or 0.1154 after Holm adjustment across seven diagnostics. I report both. The adjustment does not make the underlying concern disappear.
A wider matching caliper also made the net-payment estimate nominally significant. That means I cannot claim the significance result is stable across every specification. The dashboard lets a reader explore outcomes and trend sensitivity instead of taking a headline on trust.
12,153 matched pairs. Unadjusted nominal p = 0.1216. The adjusted interval includes zero.
| Outcome | Estimate | 95% interval |
|---|---|---|
| Unadjusted | −$0.0103 | −$0.0232 to +$0.0027 |
| Trend-adjusted | −$0.0103 | −$0.0232 to +$0.0027 |
Trend adjustment subtracts an assumed untreated difference in pre-to-post daily spending changes from the estimate and both interval endpoints, holding the standard error fixed. This is a sensitivity scenario, not an estimated correction. Calipers use matching seed 42; each may retain different customers.
Source and assumptions ↗I would not use the member spending gap to justify a rollout. I would take this as a tested analytical prototype and design a prospective experiment with contribution profit as the decision outcome.
The project taught me to ask a better business question: when customer behavior changes, how much of that change becomes value the business can actually retain?