How we measure honestly
Control group: how to honestly measure the uplift from parcel reminders
Anyone can say “our SMS raised your pickup rate”. But maybe the season, the ads or the assortment simply changed. To answer honestly, Rampo runs an A/B test of reminders: some buyers form a control group that receives no reminders from us. The uplift is the difference between the groups, and we always show it with its margin of error.
- Split by the last digit of the buyer’s phone number
- Your SMS and calls continue for both groups
- Conclusions only with enough parcels
± error
Every effect figure comes with a margin of error. Without it, the number means nothing.
What a control group is, in plain words
We split buyers into two groups by the last digit of their phone number. During the pilot, odd digits (1, 3, 5, 7, 9) form the control group and even digits form the reminder group — 50/50. After the pilot the control group narrows to digits 1 and 3, about 20% of buyers.
Both groups receive parcels on the same days, with the same advertising, in the same season. The only difference is the reminders from Rampo. So the difference in pickup rate between the groups is the effect of the service, not a calendar coincidence.
- A buyer with two orders is always in the same group
- A buyer never moves from the reminder group into control
- Control means “as you do today”, not “nothing”
How to measure SMS and reminder effectiveness: the formula
Control is your current practice
If your managers already call buyers who haven’t collected, and your CRM sends an “arrived” SMS, all of that continues for both groups. Rampo adds its own actions only in the reminder group. So we measure the uplift on top of what you already do.
How we calculate the pickup rate
For each group: pickup rate = collected ÷ (collected + not collected). Collected means received; not collected means refused or returned to sender. Parcels with an “undetermined” status are shown on a separate line and are never counted as losses.
Standard error of the difference
effect = p₁ − p₂
standard error (SE) = √( p₁(1 − p₁)/n₁ + p₂(1 − p₂)/n₂ )
Where p₁ and p₂ are the pickup rates in the reminder group and the control group, and n₁ and n₂ are the number of parcels with a known outcome in each group.
Example calculation
Reminder group: 89.4%, control group: 86.0%, 425 parcels each. Difference: +3.4 pp. The standard error is ≈ 2.2 pp, and the interval that most likely contains the true effect is about ± 4.4 pp (two standard errors). So this is a signal, not yet proof.
How many parcels you need
| Parcels per group | Standard error | A 3.5 pp effect is |
|---|---|---|
| 105 | ≈ 4.5 pp | noise |
| 425 | ≈ 2.2 pp | a signal, not proof |
| 1,050 | ≈ 1.4 pp | a visible effect |
| 2,000 | ≈ 1.0 pp | a reliable conclusion |
That is why a single small shop’s result is often within the range of chance. So we never invoice “for noise”: we show two numbers — your result with its margin of error and the combined result of all pilots. The overall “does the product work” decision is made only once each group, across all pilots, has at least 2,000 parcels with a known outcome.
Pilot timeline: 21 + 14 = 35 days
The pilot is free. Only parcels shipped within the pilot window are counted.
Day 0: start
After the offer and agreement are signed and a separate API key is connected, we switch on the call list and split buyers 50/50.
Days 1–21: shipments
Every COD parcel shipped on these days enters the pilot. Managers work the call list; the reminder group receives reminders.
Days 22–35: outcomes
Another 14 days so that every parcel is collected, returned or ends storage.
Day 35: report
Pickup rate in each group, the difference, the margin of error, “undetermined” on a separate line, and the recommended next step.
Four possible pilot outcomes
You know in advance what happens with each result. Nothing has to be decided blindly.
| What we see | What we offer | |
|---|---|---|
| Effect confirmed | Your difference is +1 pp or more | An invoice on List or Auto, only for parcels outside control |
| Not enough data | Difference within chance, no combined result yet | The pilot continues for free until the final combined result |
| Small effect for you | Your difference is within the error, but the combined result is confirmed | Pay-for-results: only for extra collected parcels |
| No effect | The combined result is below the threshold | We stop honestly, with no invoice |
We don’t promise a “guaranteed +X%”. The expected +3.5 pp uplift is a hypothesis. The control group exists precisely so that you can test it instead of taking our word for it.
Questions about the control group
Do I lose sales on the control group?
No. Control isn’t “nothing” — it is your current practice: your own SMS and manager calls continue. Those buyers just don’t get extra reminders from Rampo.
Why split by phone number rather than tracking number?
So that the same buyer with several orders is always in one group and never “leaks” between them. That keeps the measurement clean.
What is an A/B test of reminders?
A comparison of two groups of buyers that differ in one thing only — whether they receive reminders. The difference in pickup rate between them is the effect of the reminders.
Why keep a control group after the pilot?
To see every month that the effect hasn’t disappeared. After the pilot it is 20% of buyers (digits 1 and 3), and we don’t bill for their parcels.
Why is the margin of error on my result so wide?
Because the error depends on the number of parcels. With a few hundred parcels per group it is 2–4 pp. That is why we show both your result and the combined result of all pilots.
How are “undetermined” parcels handled?
On a separate line. We never count them as collected or as losses.
Test the effect on your own parcels
Start with a free audit to see your pickup rate and losses in hryvnias. Then run a 35-day pilot with a control group.