The first number was 22.3%. The number that changed the story was 40.6%.
Put members beside nonmembers after launch and the chart almost writes its own headline: members spend 22.3% more per day. It is the kind of number that invites a quick recommendation.
This was a synthetic retailer, built for a portfolio study. That gave me something an ordinary company dataset cannot offer: the ability to inspect the world I generated and compare it with a coupled world without membership benefits. But first I had to stop reading a difference between people as an effect of a program.
Before launch, the people who would eventually join were already spending 40.6% more. The apparent success story had started before there was a membership to explain it.
That did not tell me the program was useless. It told me that the first comparison was answering the wrong question. I needed a credible comparison for how members would have changed without the program.
40.6% higher spending among eventual members.
Unmatched group averages. Switching periods does not estimate a causal effect: the groups already differ before launch.
Source and assumptions ↗I found 12,153 pairs of adopters and comparable non-adopters, using observed pre-program characteristics. Then I measured the difference between their changes in daily spending.
The net-payment estimate came out at −$0.0103 per customer per day. Its interval stretched from −$0.0232 to +$0.0027. There was no clean positive revenue story to put in bold. There was also no basis for saying the true effect had to be zero.
Before discounts, estimated product value rose by $0.1367/day. After discounts and redeemed coins, the net-payment estimate was slightly negative and uncertain.
This is the finding I would bring to a decision meeting. A program can produce a positive product-value contrast without establishing higher customer payments. In the modeled accounting, discounts and rewards can absorb purchasing responses. That is consistent with these results, not a separately identified causal mechanism.
Fees were kept separate, and costs were outside the model. Calling this profit, or multiplying it into a company-wide revenue promise, would go beyond the evidence.
| Outcome | Estimate | 95% interval |
|---|---|---|
| Product value | +$0.1367 | +$0.1226 to +$0.1509 |
| Net payments | −$0.0103 | −$0.0232 to +$0.0027 |
Dots are matched DiD estimates; lines are 95% intervals. The 457-day view only rescales the same estimate, not a new forecast. Product value uses pre-discount prices; payments exclude membership fees. Neither measures profit.
Source and assumptions ↗The overall pretrend test raised a nominal warning. Widening the matching caliper to 0.10 produced p = 0.0467. Only 49.5% of adopters joined in the launch month, so the monthly event-study path mixes adoption cohorts.
None of those details belongs in a footnote that nobody can find. They shape how strongly I can interpret the estimate. I left them in the case study and exposed sensitivity controls in the dashboard.
12,153 matched pairs. Unadjusted nominal p = 0.1216. The adjusted interval includes zero.
| Outcome | Estimate | 95% interval |
|---|---|---|
| Unadjusted | −$0.0103 | −$0.0232 to +$0.0027 |
| Trend-adjusted | −$0.0103 | −$0.0232 to +$0.0027 |
Trend adjustment subtracts an assumed untreated difference in pre-to-post daily spending changes from the estimate and both interval endpoints, holding the standard error fixed. This is a sensitivity scenario, not an estimated correction. Calipers use matching seed 42; each may retain different customers.
Source and assumptions ↗The coupled counterfactual put the known net-payment ATT at about −$0.01128/day, inside the reported interval. That supports the implementation in this simulated world. It does not certify the assumptions in a real company.
I also ran 40 smaller simulation replications, kept the uncertainty around that validation visible, and built a pipeline that checks whether the reports agree with the data.
The ending is a decision, not a victory lap: test contribution profit prospectively before recommending a real rollout. The useful finding is the separation between buying more product value and paying more money.